Skip to main content
Glama

Server Details

100+ agent-payable C-suite expertises with x402 micro-payments — competitive intel, SEC filings, sanctions, KYC, clinical evidence, real estate, ESG. 183 tools, free tier 100 calls/month.

Status
Healthy
Last Tested
Transport
Streamable HTTP
URL

Glama MCP Gateway

Connect through Glama MCP Gateway for full control over tool access and complete visibility into every call.

MCP client
Glama
MCP server

Full call logging

Every tool call is logged with complete inputs and outputs, so you can debug issues and audit what your agents are doing.

Tool access control

Enable or disable individual tools per connector, so you decide what your agents can and cannot do.

Managed credentials

Glama handles OAuth flows, token storage, and automatic rotation, so credentials never expire on your clients.

Usage analytics

See which tools your agents call, how often, and when, so you can understand usage patterns and catch anomalies.

100% free. Your data is private.
Tool DescriptionsC

Average 3.6/5 across 271 of 271 tools scored. Lowest: 1.8/5.

Server CoherenceD
Disambiguation2/5

With 271 tools, many have overlapping purposes (e.g., multiple competitor intel tools, multiple financial modelers, multiple ESG auditors). Detailed descriptions help slightly, but the sheer volume creates confusion. Agents would struggle to select the right tool among many similar options.

Naming Consistency1/5

Tool names are wildly inconsistent: mix of English and French, snake_case and short phrases, some very generic (process, run, execute equivalents). No discernible naming convention (e.g., abm_architect vs. boundary_control vs. bp_narratif). This makes it hard to predict tool names.

Tool Count1/5

271 tools is far beyond typical well-scoped servers (3-15). This indicates an unfocused, over-bloated tool surface. Even for a general business intelligence server, this number is excessive and violates the principle of each tool earning its place.

Completeness2/5

Despite the large count, coverage feels scattered. Some domains (e.g., content, competitive intel) have many tools, while others (e.g., supply chain, HR) have gaps. The set lacks a coherent scope; it seems like a dump of many separate tool collections rather than a complete, curated surface.

Available Tools

279 tools
abm_architectC
Read-only
Inspect

Architecte ABM — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Gapup Hub — ABM 20 comptes nommés · Budget €120k · Tier 1×5 + Tier 2×15 · Playbooks 3 niveaux. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
productYes
salesTeamNo
icpCriteriaYes
abmBudgetEurNo
targetAccountsYes
currentChannelsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavioral context: 'Returns a structured, audited deliverable' and 'Inputs are validated server-side.' Annotations already declare readOnlyHint and openWorldHint, so the description complements these without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat concise but mixes languages and includes a specific reference case that may not be universally relevant. It could be more focused and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, nested objects, no output schema), the description is insufficient. It does not explain the output format beyond 'structured, audited deliverable,' leaving the agent without enough information to correctly invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (13%). The description does not elaborate on any parameter meanings, merely stating to 'send the documented case fields.' It fails to compensate for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Architecte ABM' and 'Returns a structured, audited deliverable,' indicating the tool generates an ABM deliverable. However, the purpose is vague due to marketing language and lack of a clear verb+resource statement. It does not differentiate from siblings like 'ld_architect' or 'abm_lookalike_account_finder'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description only mentions to 'send the documented case fields,' but does not explain context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

abm_lookalike_account_finderA
Read-onlyIdempotent
Inspect

As a CMO, discover 50 B2B accounts that closely match your top 10 customers' tech stacks and firmographics. This tool analyzes public web data including robots.txt and OpenGraph metadata to identify lookalike accounts for targeted ABM campaigns. Input your top customer domains and desired firmographic filters to receive a ranked list of potential targets with matching technologies and company attributes.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
tech_stack_keywordsNoSpecific technologies to match in lookalike accounts
firmographic_filtersNo
top_customer_domainsYesList of top 10 customer domains to use as seed accounts

Output Schema

ParametersJSON Schema
NameRequiredDescription
statsNo
statusYes
sourcesYes
warningsYes
lookalike_accountsYes
matched_technologiesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint. The description adds context about data sources (robots.txt, OpenGraph metadata) and mentions output ranking. No contradictions, but could elaborate on async behavior and rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each adding value: action, method, inputs/outputs. Front-loaded with the core purpose. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, description adequately covers usage. It explains the tool's function and data sources, though it could mention the async parameter as an operational detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%. Description adds meaning by naming the firmographic filters and stating the output count (50 accounts). It complements the schema without repeating it fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as discovering 50 B2B lookalike accounts based on tech stacks and firmographics. It specifies the input (top 10 customer domains) and output (ranked list), distinguishing it from siblings like account_expansion_mapper.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (for targeted ABM campaigns) and what inputs are needed. However, it does not explicitly mention when not to use it or compare with alternatives among the sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

account_expansion_mapperB
Read-only
Inspect

Mapping d'expansion comptes — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Notion B2B Enterprise — top 30 strategic accounts · expansion plays NRR 130%+ target · Snowflake/Shopify/Vercel/Stripe analyzed. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
accountsYes
ownershipYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint, and the description adds that inputs are validated server-side and output is an audited deliverable. This provides useful behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with four sentences, front-loading the purpose. However, it mixes French and English, which slightly reduces clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested schema and no output schema, the description is incomplete. It mentions a deliverable but gives no detail on its structure or content, and the reference case is not comprehensive enough to guide usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (20%), and the description adds no explanation of parameters. It only says 'send the documented case fields', which does not clarify the specific fields or their meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it performs expansion account mapping for CRO and returns a structured deliverable. It is clear but does not differentiate from siblings like 'abm_architect' or 'upsell_hunter', which are also about account growth.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description includes a reference case but does not specify prerequisites or when to choose this over similar tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

action_plan_esgC
Read-only
Inspect

Plan d'action ESG — Gapup agent-payable C-suite expertise (SUSTAINABILITY). Returns a structured, audited deliverable. Reference case: TechCorp SAS — Plan ESG 36 mois (500 FTE, €60M CA, score 54→76/100). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
horizonYes36 mois
ambitionsYes
targetLabelsNo
currentScoresNo
availableResourcesYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint: true, openWorldHint: true) already indicate a read-only, open-world operation. The description adds only that it returns a 'structured, audited deliverable' and mentions a reference case. No behavioral details like auth, rate limits, or side effects beyond the annotation-provided traits. Minimal added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise but includes an extensive reference case ('TechCorp SAS...') that may be extraneous. The key information is front-loaded, but the example could be more succinct without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema is provided, and the description only says 'structured, audited deliverable' without detailing the return format, fields, or how to interpret results. For a tool with 8 parameters and nested objects, this is insufficient for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13%, leaving 87% undocumented. The description provides no parameter explanations, merely stating 'Inputs are validated server-side'. It fails to add meaning to any of the eight parameters, including nested objects like company, ambitions, etc.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates an ESG action plan and returns a structured deliverable. It provides a reference case and mentions 'agent-payable C-suite expertise', giving specificity. However, it does not explicitly differentiate from siblings like esg_audit_multi or sustainability_report, which are ESG-related.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool over alternatives. The description only says 'send the documented case fields' but lacks context on prerequisites, typical use cases, or exclusions. With many ESG siblings, this is a significant gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

adversarial_input_stress_testerA
Read-onlyIdempotent
Inspect

An asynchronous risk assessment tool that evaluates AI model resilience against adversarial inputs following NIST AI Risk Management Framework (RMF) red-teaming protocols. Designed for security and compliance personas, it accepts model outputs or decision boundaries and returns structured risk scores, failure modes, and adversarial examples. Requires async:true to avoid timeout errors. Outputs include status, warnings, and source references.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
maxTestsNoMaximum number of adversarial tests to run
modelOutputYesThe AI model's output or decision to be stress-tested
adversarialDatasetNoOptional custom adversarial inputs to test
sensitivityThresholdNoThreshold for flagging high-risk adversarial examples

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
riskScoreNoNormalized risk score from adversarial testing
failureModesNo
adversarialExamplesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds behavioral details: it is asynchronous, requires the async flag, and outputs status, warnings, and source references. It aligns with annotations (no contradiction) and adds context about timeout avoidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with purpose, then persona, async requirement, and outputs. Every sentence adds value without redundancy or verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 params, async, NIST RMF), the description covers purpose, usage hint, and output types. The existence of an output schema (has output schema: true) reduces the burden. It could elaborate on NIST RMF protocols, but overall it is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers all 5 parameters with descriptions (100% coverage). The description repeats the modelOutput parameter ('accepts model outputs or decision boundaries') and mentions the async flag requirement, but adds little beyond schema for maxTests, adversarialDataset, and sensitivityThreshold. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates AI model resilience against adversarial inputs following NIST AI RMF protocols, specifying it accepts model outputs or decision boundaries and returns structured risk scores, failure modes, and adversarial examples. This distinguishes it from siblings like jailbreak_attempt_detector or safety_guardrail_breach_analyzer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly notes 'Requires async:true to avoid timeout errors', guiding usage for performance. It also mentions it is 'designed for security and compliance personas', providing context. However, it lacks explicit when-not-to-use or alternative tool references, though siblings exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

affiliate_fraud_clickstream_detectorA
Read-onlyIdempotent
Inspect

Analyzes affiliate clickstream data from Common Crawl to flag potential fraud patterns (duplicate IPs, rapid clicks, device spoofing). Designed for CMOs to validate affiliate traffic quality and prevent budget waste. Inputs: affiliate network name and date range. Outputs: fraud probability score, suspicious IP list, and pattern analysis. Keywords: affiliate fraud detection, clickstream analysis, marketing attribution, traffic validation.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
thresholdNoFraud probability threshold (0.1-0.99)
date_rangeYes
affiliate_networkYesName of the affiliate network to analyze (e.g., 'CJ Affiliate', 'Rakuten')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
suspicious_ipsNo
fraud_probabilityNoOverall fraud probability score (0-1)
patterns_detectedNo
total_clicks_analyzedNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds value by detailing the output structure (fraud probability score, suspicious IP list, pattern analysis) and input requirements (affiliate network, date range). It is consistent with annotations and does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences plus a keywords tag, each serving a purpose: purpose, audience/context, inputs/outputs. It is efficient and front-loaded with the core action. However, the keywords tag is redundant and could be omitted for even greater conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, nested objects, and an output schema (as per context signals), the description adequately covers purpose, inputs, and outputs. It mentions the data source (Common Crawl) and target audience. It does not discuss edge cases or failure modes, but this is acceptable given the read-only, idempotent nature.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, and the input schema already documents all parameters (affiliate_network, date_range, threshold, async). The description merely restates that inputs are 'affiliate network name and date range', adding no new meaning beyond the schema for these or the other parameters. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Analyzes') and resource ('affiliate clickstream data from Common Crawl') and clearly states the goal ('flag potential fraud patterns'). It lists example fraud types (duplicate IPs, rapid clicks, device spoofing), distinguishing it from generic fraud detectors. The target audience (CMOs) and business value (prevent budget waste) are also provided.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context ('Designed for CMOs to validate affiliate traffic quality') but does not explicitly state when not to use the tool or compare it to sibling tools like 'fraud_detector' or 'web_search_multilang'. It lacks explicit exclusions or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

africa_trade_barrier_breakerA
Read-onlyIdempotent
Inspect

As a COO, analyze non-tariff trade barriers (NTBs) across African trade corridors using WITS and UNCTAD STAT data. Input origin/destination countries and product HS codes to receive barrier mapping with severity scores and actionable mitigation strategies. Returns structured risk assessment, regulatory compliance gaps, and supply chain optimization recommendations. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
hs_codeNo6-digit Harmonized System product code
origin_countryYesISO 3-letter country code for export origin
destination_countryYesISO 3-letter country code for import destination
include_regulatory_detailsNoWhether to include detailed regulatory text in output

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
warningsYes
barrier_summaryYes
trade_flow_impactNo
regulatory_detailsNo
mitigation_strategiesYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint, openWorldHint, and idempotentHint. The description adds behavioral context by mentioning that passing async:true avoids timeout and describing the return format (structured risk assessment, etc.). This supplements the annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, front-loading the core function, then detailing inputs and outputs. It is concise with no wasted words, and the async instruction is placed at the end for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given complexity, the description covers key aspects: data sources, inputs, outputs, and async behavior. An output schema exists (though not shown), so the description doesn't need to detail returns. Slightly lacking in explaining severity scoring or mitigation strategy format, but overall sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 5 parameters are fully described in the input schema (100% coverage). The description adds minimal additional meaning beyond highlighting origin/destination countries and HS codes. The async parameter usage note is helpful but not critical. Baseline 3 applies as schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: analyzing non-tariff trade barriers across African trade corridors using specific data sources. It mentions the role (COO), inputs (origin/destination countries, HS codes), and outputs (barrier mapping with severity scores, mitigation strategies). This differentiates it from sibling tools that focus on other Africa-related trade analyses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by a COO for analyzing NTBs, but it does not explicitly state when to use this tool versus alternatives. It lacks guidance on when not to use it or how it compares to sibling tools like africa_trade_finance_esg_rater or africa_trade_preference_arbitrage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

africa_trade_finance_esg_raterA
Read-onlyIdempotent
Inspect

As a COO, evaluate ESG compliance of African trade finance providers using World Bank WITS trade statistics and CDP climate disclosure data. Input the financial institution's name or identifier, and receive an ESG rating with breakdown across environmental, social, and governance dimensions. Ideal for due diligence on trade partners or portfolio risk assessment. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNoAssessment year (2018-2023)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
countryCodeNoISO 2-letter country code (e.g., 'ZA' for South Africa)
institutionNameYesFull name of the trade finance provider (e.g., 'Standard Bank Group')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
warningsYes
esgRatingYes
socialScoreNo
tradeVolumeNoAnnual trade finance volume (USD)
carbonIntensityNoCO2 emissions per million USD financed (tons)
governanceScoreNo
environmentalScoreNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnly, openWorld, and idempotent hints. The description adds value by noting the 'async' parameter to avoid timeouts, which is a behavioral trait not covered by annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise at two sentences plus an async note, but it could be better structured with bullet points or clearer separation of purpose vs. usage. Still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema and 4 parameters, the description provides context on data sources and use cases. It adequately covers the core function, though more detail on output structure or parameter relationships could help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, providing baseline of 3. The description adds meaning to the 'async' parameter with usage guidance and reiterates the primary input (institution name), improving understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates ESG compliance of African trade finance providers using specific data sources (World Bank WITS, CDP) and returns an ESG rating with dimensional breakdown. It distinguishes itself from siblings by its geographic and sector focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions ideal use cases ('due diligence on trade partners or portfolio risk assessment') but does not explicitly state when to avoid this tool or mention alternative tools for similar tasks, leaving ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

africa_trade_preference_arbitrageA
Read-onlyIdempotent
Inspect

Analyzes AGOA (African Growth and Opportunity Act) and EBA (Everything But Arms) trade preference arbitrage opportunities for COOs evaluating export strategies. Compares tariff rates, trade volumes, and preference utilization across eligible African countries using WITS and OECD trade data. Returns structured analysis of potential duty savings, market access advantages, and compliance requirements. — pass async:true REQUIRED to avoid x402 timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNoReference year for trade data
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
hs_codeYes6-10 digit Harmonized System product code
exporting_countryYesISO 2-letter country code of African exporter
importing_countryNoISO 2-letter country code of target market (US/EU)
preference_schemeNoTrade preference scheme to analyze

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
duty_savings_pctNoEstimated duty savings percentage under preference scheme
trade_volume_usdNoAnnual trade volume in USD for given HS code
market_access_scoreNoComposite score of market access advantage (0-100)
compliance_requirementsNoList of compliance requirements for preference eligibility
preference_utilization_rateNoPercentage of eligible exports utilizing preference
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint. The description adds valuable behavioral context: it uses WITS and OECD data, returns structured analysis, and crucially warns about the async parameter to avoid timeouts. This exceeds annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences with no redundancy: purpose, details, output, critical usage note. Could be slightly more structured (e.g., bullet points) but efficient overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters and an output schema, the description covers purpose, data sources, output nature, and a critical usage note. It doesn't explain the output schema, but since one exists, the burden is lower. Largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all 6 parameters. The description adds no extra meaning beyond what schema provides, except the async timeout note which is already in the parameter description. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes AGOA/EBA trade preference arbitrage opportunities for COOs, comparing tariff rates, trade volumes, etc. It distinguishes from siblings like 'africa_trade_preference_optimizer' only implicitly; explicit differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description targets 'COOs evaluating export strategies' but provides no guidance on when to use this tool versus alternatives (e.g., 'agoa_eba_intelligence' or 'tariff_arbitrage_finder'). No exclusions or context for siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

africa_trade_preference_optimizerA
Read-onlyIdempotent
Inspect

As a COO, analyze AGOA/EBA duty savings opportunities with HS code-level trade route optimization. Input origin country, destination country, and HS code to receive duty savings estimates, optimal trade routes, and preference utilization recommendations. Uses UN Comtrade trade flow data, WCO tariff schedules, and African Union trade agreement rules. Ideal for export market evaluation, supply chain optimization, and trade agreement compliance analysis. Keywords: AGOA, EBA, duty savings, trade optimization, HS code, African trade, export strategy.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
hsCodeYes6-10 digit Harmonized System code (e.g., '010121' for live horses)
quantityNoEstimated annual export quantity in units
valueUsdNoEstimated annual export value in USD
originCountryYesISO 3166-1 alpha-3 country code of export origin (e.g., 'KEN' for Kenya)
destinationCountryYesISO 3166-1 alpha-3 country code of import destination (e.g., 'USA' for United States)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
dutySavingsNoEstimated annual duty savings in USD under optimal preference program
optimalRouteNo
alternativeRoutesNo
complianceWarningsNoPotential compliance risks or documentation requirements
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, openWorldHint, and idempotentHint. The description adds useful behavioral context, such as using external data sources (UN Comtrade, WCO tariff schedules, African Union rules) and generating specific outputs (duty savings estimates, optimal trade routes, recommendations). This goes beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with the main action in the first sentence. It lists inputs, outputs, data sources, and use cases efficiently. The inclusion of keywords at the end is slightly redundant but does not detract significantly from conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and annotations, the description provides adequate context: it explains the tool's purpose, inputs, outputs, and data sources. However, it does not mention the async behavior of the optional parameter, and there is no discussion of error conditions or data freshness, which would improve completeness for a data-dependent tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so all parameters are documented in the schema. The description mentions the three required parameters (originCountry, destinationCountry, hsCode) but does not add extra semantics beyond what the schema provides for optional parameters like async, quantity, and valueUsd. Therefore, the description adds minimal value in this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool analyzes AGOA/EBA duty savings opportunities with HS code-level trade route optimization, specifying inputs and outputs. However, it does not explicitly differentiate from sibling tools like africa_trade_preference_arbitrage or agoa_eba_intelligence, which may have overlapping purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions it is 'Ideal for export market evaluation, supply chain optimization, and trade agreement compliance analysis,' providing context for when to use the tool. However, it does not include explicit guidance on when not to use it or any comparisons to alternative tools, which is a gap given the many sibling tools in the same domain.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agoa_eba_intelligenceA
Read-only
Inspect

Intelligence préférentielle AGOA (US→Africa) et EBA/GSP (EU→Africa). Vérifie l'éligibilité d'un pays africain aux programmes tarifaires préférentiels, l'éligibilité d'un produit par code HS, identifie les meilleures opportunités d'export Afrique→US/EU, et fournit les règles de conformité (rules of origin, valeur ajoutée, docs). Différenciateur Africa diaspora : 39 pays AGOA + 47 LDCs EBA encodés. Sources : AGOA.info · EU EBA · EU GSP+ · WTO Tariff · UN Comtrade.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesMode d'analyse : 'country_eligibility' (statut AGOA/EBA/GSP d'un pays africain) | 'product_eligibility' (éligibilité d'un produit par code HS) | 'trade_opportunity' (top opportunités export Afrique→US/EU) | 'compliance_check' (rules of origin, seuils valeur ajoutée, documentation)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
hs_codeNoCode HS (Harmonized System) 6+ chiffres (requis pour product_eligibility). Exemple : '620342' = pantalons coton homme, '090111' = café arabica non torréfié, '060310' = fleurs fraîches.
country_isoNoCode ISO 2-lettres du pays africain (requis pour country_eligibility). Exemples : KE=Kenya, NG=Nigeria, ZA=Afrique du Sud, ET=Éthiopie, LS=Lesotho, GH=Ghana.
destinationNoMarché de destination pour trade_opportunity : 'US', 'EU', ou 'both' (défaut). Ignoré pour les autres modes.
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the description does not need to restate safety. The description adds context about the tool's coverage (countries, sources) but does not disclose other behavioral traits like synchronous/asynchronous behavior or response format. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that efficiently conveys the tool's purpose and capabilities. It is front-loaded with the core function. While it includes some redundancy with the parameter descriptions, it remains concise for a tool with multiple modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the tool's complexity (5 parameters, 4 modes), the description does not explain the output format or what the tool returns. Since there is no output schema, this is a significant gap. It also lacks information on error handling or response behavior. The description is incomplete for an agent to fully understand the tool's results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All parameters are documented in the schema with descriptions (100% coverage). The tool description provides an overview but adds minimal additional meaning beyond listing the modes. The schema already explains mode options, hs_code examples, etc., so the description does not significantly enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: checking eligibility for AGOA/EBA/GSP programs, product eligibility by HS code, identifying trade opportunities, and providing compliance rules. It distinguishes itself from siblings by mentioning the specific scope of 39 AGOA countries and 47 LDC EBA countries, making it unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through its list of modes (country_eligibility, product_eligibility, etc.), but does not explicitly state when to use this tool versus related siblings like africa_trade_preference_arbitrage or africa_trade_preference_optimizer. No guidance on prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ai_act_incident_responseA
Read-onlyIdempotent
Inspect

Generates EU AI Act incident response playbooks with regulator notification templates for risk management teams. Inputs include incident severity, AI system type, and affected stakeholders. Outputs structured playbook steps, regulator notification drafts, and compliance checklists. Essential for high-risk AI system breaches requiring formal EU notification — pass async:true REQUIRED to avoid x402 timeout. Keywords: AI Act compliance, incident response, regulator notification, risk management, ISO 27035, NIST SP 800-61.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
severityYes
incident_typeYes
ai_system_typeNo
incident_descriptionNo
affected_stakeholdersNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
next_stepsNo
playbook_stepsNo
compliance_checklistNo
regulator_notificationNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds value beyond annotations by detailing the async requirement and the types of outputs (playbook steps, notification drafts, compliance checklists). However, there is a slight contradiction with the readOnlyHint annotation, as 'generates' implies creation, but the annotation claims the tool is read-only. This reduces transparency slightly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, starting with the tool's purpose and followed by inputs, outputs, and a critical usage note about async. The keywords at the end are slightly redundant but not detrimental.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description does not need to explain return values. It covers the core purpose, inputs, and outputs, along with the async requirement. It could be more explicit about the playbook structure or compliance standards, but overall it provides sufficient context for an agent to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 17% schema description coverage, the description partially compensates by listing three parameters (incident severity, AI system type, affected stakeholders) and their role in generating outputs. However, it does not cover all parameters (e.g., async, incident_description) or provide format details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates EU AI Act incident response playbooks with regulator notification templates, specifying inputs and outputs. However, it does not explicitly differentiate from sibling tools like ai_act_sandbox_regulatory_sandbox or incident_response_evidence_collector, leaving some ambiguity about when to use this specific tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for high-risk AI system breaches requiring formal EU notification and advises passing async:true to avoid timeout, but it does not explicitly state when not to use this tool or mention alternatives. The guidance is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ai_act_sandbox_regulatory_sandboxA
Read-onlyIdempotent
Inspect

A legal-focused tool for simulating EU AI Act regulatory sandbox submissions. Provides structured feedback on compliance, risk levels, and required documentation based on EUR-Lex and OECD AI Policy Observatory sources. Accepts AI system descriptions, intended use cases, and technical specifications as input. Returns detailed assessment with warnings, citations, and actionable recommendations for legal teams and AI developers.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
sectorNoPrimary sector of application
riskLevelYesSelf-assessed risk level of the AI system
intendedUseYesPrimary and secondary use cases of the AI system
documentationNoList of provided documentation types (e.g., 'technical', 'ethical', 'data')
systemDescriptionYesDetailed description of the AI system including purpose, architecture, and data sources

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
assessmentNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, open-world, and idempotent behavior. The description adds value by detailing the nature of feedback (warnings, citations, recommendations) and the sources used (EUR-Lex, OECD). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, each contributing essential information. It is front-loaded with the core purpose and efficiently covers inputs, outputs, and target audience without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 params, output schema exists), the description adequately covers the tool's purpose, inputs, and output nature. The return value is described ('detailed assessment with warnings, citations, and actionable recommendations'), and the output schema addresses formal structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents parameters well. The description provides general context (e.g., 'technical specifications') but does not add significant meaning beyond what the parameter descriptions offer. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: simulating EU AI Act regulatory sandbox submissions. The verb 'simulating' and resource 'regulatory sandbox submissions' are specific. It distinguishes from sibling AI Act tools (e.g., ai_act_incident_response) by focusing on sandbox simulation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for legal teams and AI developers preparing sandbox submissions, but it does not explicitly state when to use this tool versus related alternatives (e.g., ai_act_incident_response). No exclusions or when-not guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ai_act_training_data_auditA
Read-onlyIdempotent
Inspect

As a CTO, audit AI training datasets for EU AI Act compliance with bias detection and regulatory risk assessment. Inputs: dataset identifier (Hugging Face ID or URL) and optional risk thresholds. Outputs: compliance score, bias metrics, regulatory warnings, and source references. Ideal for pre-deployment risk evaluation. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
dataset_idYesHugging Face dataset identifier or direct URL to dataset
risk_thresholdNo
include_bias_metricsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
bias_metricsNo
compliance_scoreNo
dataset_metadataNo
regulatory_warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, indicating a safe, read-only, idempotent operation. The description adds extra context: it advises passing async:true to avoid timeout, implying the tool can be long-running. It also specifies the return includes a job_id when async is true. The description does not contradict annotations and provides useful behavioral guidance beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences. The first sentence defines purpose and outputs. The second lists inputs and outputs. The third gives use case and a crucial tip about async. Every sentence earns its place with no fluff or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description, combined with the full input schema and annotations (readOnlyHint, openWorldHint, idempotentHint), provides a complete picture. The tool has an output schema (not shown but confirmed), so return values are documented. The description covers inputs, outputs, use case, and async behavior. It comprehensively sets expectations for a pre-deployment audit tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides descriptions for all four parameters (dataset_id, risk_threshold, include_bias_metrics, async). The tool description adds minimal value: it specifies that dataset_id is a Hugging Face ID or URL (already in schema) and that risk_threshold is optional (also implied by default). The async parameter is mentioned in context, but the schema already explains its behavior. With high schema coverage, the description's contribution is limited.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool audits AI training datasets for EU AI Act compliance with bias detection and risk assessment. It specifies exact inputs (dataset identifier and optional risk thresholds) and outputs (compliance score, bias metrics, etc.). This distinguishes it from sibling tools like ai_act_incident_response or ai_act_sandbox_regulatory_sandbox, which focus on other aspects of AI Act compliance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Ideal for pre-deployment risk evaluation' and advises using async:true to avoid timeout. It provides context on when to use (before deployment) but does not explicitly say when not to use or compare to alternatives like ai_governance_pilot or the full report tools. A clear use case is given, but exclusions are missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ai_governance_full_report_asyncA
Read-only
Inspect

Audit EU AI Act complet (Règlement UE 2024/1689) — implémentation native audit-grade. Classifie le système IA selon les 4 tiers de risque (unacceptable/high_risk/limited_risk/minimal_risk/gpai) sur la base de l'Annexe III et de l'Article 5. Produit : (1) classification tier + justification + articles applicables, (2) checklist conformité Articles 9-15 + 50 + 53-55, (3) gaps documentation Annexe IV, (4) mapping ISO 42001, (5) deadlines EU AI Act 2025-2029, (6) estimation coût et effort, (7) top 10 recommandations P0/P1/P2. Retourne immédiatement (<300ms) un job_id. Poller avec ai_governance_full_report_result(job_id) après eta_seconds (~90s). Cache 7 jours pour inputs identiques. Async tool — register a webhook via webhooks_manage(register, url, [job.completed]) to receive callbacks instead of polling. Faster + lighter. DISCLAIMER : non substitutif à un avis juridique professionnel.

ParametersJSON Schema
NameRequiredDescriptionDefault
company_sizeNoTaille entreprise : startup (≤50), smb (51-250), mid (251-1000), large (1001-5000), enterprise (>5000)
data_sourcesNoSources de données utilisées par le système IA
affected_personsNoCatégories de personnes affectées par les décisions du système (ex: candidats, employés, clients)
geographic_scopeNoZones géographiques de déploiement (ex: 'EU', 'France', 'Global')
intended_purposeYesFinalité prévue du système IA : à quoi sert-il concrètement
deployment_contextNoContexte de déploiement : interne (usage employés), public, B2B, B2C
ai_system_descriptionYesDescription détaillée du système IA : ce qu'il fait, comment il fonctionne, quelles décisions il prend

Output Schema

ParametersJSON Schema
NameRequiredDescription
job_idYesIdentifiant unique du job — passer à ai_governance_full_report_result
statusYes
eta_secondsYesDurée estimée avant disponibilité du résultat
submitted_atYesTimestamp ISO-8601 de soumission
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses async nature (<300ms return), caching duration (7 days), and disclaimer. No contradiction with annotations; readOnlyHint aligns with non-mutating audit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with bullets but slightly verbose. Could be more concise while retaining key info.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters and output schema, description thoroughly explains async workflow, caching, webhooks, and disclaimer. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions; description adds no extra parameter detail. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it audits EU AI Act, classifies risk tiers, and lists 7 specific outputs. It distinguishes from sibling tool ai_governance_full_report_result by noting async polling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains when to use: submit async and poll with result tool or register webhook. Provides alternatives (polling vs callback) and mentions caching behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ai_governance_full_report_resultA
Read-onlyIdempotent
Inspect

Poll the result of an ai_governance_full_report_async job. Returns status=pending while running, status=completed with the full EU AI Act governance audit report once done (risk_tier, compliance checklist Articles 9-15/50/53-55, Annex IV documentation gaps, ISO 42001 alignment, deadlines 2025-2029, cost estimate, top-10 recommendations P0/P1/P2, compliance_score), status=failed on error, or status=not_found if the job_id is unknown or expired (TTL 24h). Call this after the eta_seconds hint returned by ai_governance_full_report_async (~90s).

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by ai_governance_full_report_async (prefix: aigfr_)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description discloses all statuses (pending, completed, failed, not_found), TTL 24h, and details of the report content. Annotations already indicate read-only and idempotent; description adds significant behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single paragraph containing all necessary information without redundancy. Could be slightly more structured but is efficient and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of polling an async job with multiple statuses and a detailed report, the description covers all aspects: statuses, TTL, usage timing, and report contents. Output schema exists, so return values are documented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (job_id) with schema description covering prefix constraint. Schema coverage is 100%, so description adds minimal extra value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states verb 'Poll' and resource 'result of an ai_governance_full_report_async job'. Distinguishes from sibling tools by specifying it's for polling results after async initiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call: after the eta_seconds hint (~90s). Does not explicitly mention when not to use, but context implies it should only be called after the async job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ai_governance_pilotC
Read-only
Inspect

Pilotage de gouvernance IA — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: TalentScope SAS — scoring IA candidats RH (EU AI Act Annex III §4, high-risk). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
aiUseCasesYes
targetFrameworksYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. The description adds that inputs are validated server-side and that the tool returns a structured deliverable, which is consistent. It does not disclose potential costs, rate limits, or details about the deliverable's format, but the added context is moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise with four sentences, but it includes unnecessary jargon ('Gapup agent-payable C-suite expertise (RISK)') that reduces clarity. The key information is front-loaded, but the structure could be more efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has complex nested inputs (5 parameters, 3 required) and no output schema. The description only specifies a 'structured, audited deliverable' without detailing its content or response format. For the complexity, the description is insufficient for an agent to correctly invoke the tool and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, so the description should compensate, but it fails to explain parameters beyond mentioning 'send the documented case fields'. The reference case hints at what inputs like company, aiUseCases, and targetFrameworks mean, but it is vague and in French. The async parameter and focus parameter are ignored.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool pilots AI governance and returns a structured, audited deliverable. It includes a reference case that clarifies the domain (high-risk AI use in HR). However, it does not distinguish this tool from siblings like ai_governance_full_report_async or vertical_ai_agent_governance, which may have overlapping purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives. It does not mention prerequisites, limitations, or comparison to sibling tools. The openWorldHint annotation suggests flexibility, but the description lacks actionable usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

anti_demissions_hrC
Read-only
Inspect

Bouclier anti-démissions — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: Buffer Inc — détection des at-risk parmi 80 FTEs (Q1 2026). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
signalsYes
employeesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. Description adds that inputs are validated server-side, but no further behavioral traits beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is short but includes jargon (Gapup agent-payable C-suite expertise) that may confuse. Structure is acceptable but not optimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, so description should explain return value. It only mentions 'structured, audited deliverable' without details on format or interpretation. Complex nested inputs are not compensated for.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (20%). Description does not explain any parameters, leaving the agent without additional meaning for the complex nested inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes employee data to detect attrition risk and returns a structured deliverable. However, it does not differentiate from siblings like churn_defender or talent_poaching_risk.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. No explicit context or exclusions provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arbitration_awards_lookupA
Read-onlyIdempotent
Inspect

Commercial arbitration intelligence for litigation lawyers, M&A due diligence teams, sovereign wealth funds and trade finance compliance. Covers 8 major institutions: ICC, AAA, LCIA, HKIAC, SIAC, CIETAC, DIAC, ICDR.

Three modes: • party_lookup — find awards by party name (searches 20 landmark public awards + JusMundi best-effort) • institution_index — browse awards and caseload stats per institution with date range filter • clause_check — audit an arbitration clause for missing elements (institution, seat, language, arbitrator count, governing law, binding nature)

Note: Most arbitration awards are confidential. This tool surfaces public awards (Yukos, Crystallex, Achmea, etc.) plus redacted statistics from institutional annual reports. Private awards are not accessible.

Cache: 24h (arbitration data is very stable). No API key required.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesparty_lookup: search by party name or keyword. institution_index: browse awards by institution + stats. clause_check: audit an arbitration clause for issues.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesFor party_lookup: party name or keyword (e.g. "Yukos", "Russia"). For institution_index: institution name or keyword. For clause_check: full text of the arbitration clause to audit.
date_toNoISO date filter to (YYYY-MM-DD). Applied to award_date.
date_fromNoISO date filter from (YYYY-MM-DD). Applied to award_date.
institutionNoFilter by institution. Default 'all'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
queryYes
awardsNo
statusYes
sourcesYes
clause_checkNo
quality_scoreYes
institution_statsNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, and non-destructive behavior. The description adds valuable context: confidentiality of awards, 24-hour cache stability, and no API key requirement. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections: audience, covered institutions, three modes, confidentiality note, and caching info. It is front-loaded with essential purpose and concise with no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, 3 modes, 8 institutions, confidentiality constraints) and the presence of an output schema, the description covers all necessary context: modes, institutions, confidentiality, caching, and authentication.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds mode-specific guidance for the 'query' parameter and explains how each mode uses it, providing additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: commercial arbitration intelligence with three distinct modes (party_lookup, institution_index, clause_check). It specifies the target audience and covers 8 major institutions, distinguishing it from sibling tools like 'legal_clause_extractor'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implicit guidance by explaining that most awards are confidential and that the tool only surfaces public awards and redacted statistics. It mentions caching and lack of API key, but does not explicitly compare to alternative tools for when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attack_surface_monitorB
Read-only
Inspect

Surveillance surface d'attaque — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Answers: Which Internet-facing assets of combine a critical CVE, an exposed service, and no WAF — top findings to fix in 14 days? · What is the attack surface of : subdomains, open ports, SSL/TLS grades, and associated CVEs? · Give me a CISO-ready ASM report with blast radius estimate and SLA-driven remediation plan for . · What is the email phishing risk for ? Assess SPF/DMARC posture and recommend improvements. · During M&A due diligence, what are the top cyber exposures on 's Internet-facing infrastructure? Reference case: Velora Payments — 8 assets exposés · 2 critiques (CVE-2023-44487 HTTP/2 RapidReset, Admin panel ouvert) · . Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
domainYes
exclusionsNo
scope_cidrsNo
include_email_surfaceYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations readOnlyHint=true and openWorldHint=true are consistent with the description, which says it returns a report and mentions external domain analysis. The description adds context that the deliverable is 'audited' and mentions a 14-day fix suggestion, but does not cover rate limits, authentication needs, or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is overly long, contains marketing fluff ('Gapup agent-payable C-suite expertise (RISK)'), and lacks a concise, front-loaded summary. It includes multiple example questions that could be streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 6 parameters (2 required) and no output schema. The description gives examples of outputs but does not explain the structured deliverable's format, handling of async, or how parameters like 'focus' and 'exclusions' affect results. It is insufficient for an agent to fully understand the tool's capabilities.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only 'async' has a description). The tool description does not explain the meaning or usage of parameters like 'focus', 'exclusions', 'scope_cidrs', or 'include_email_surface' beyond implying email surface through an example. This fails to compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it performs attack surface monitoring for a domain, returning a structured audit deliverable. It gives clear examples of what it can do (find critical CVEs, email phishing risk, etc.). However, it does not differentiate from similar sibling tools like 'cve_security_lookup' or 'cyber_risk_auditor'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit example queries that serve as usage guidelines, such as 'Which Internet-facing assets of <domain> combine a critical CVE...' and 'During M&A due diligence...'. It does not explicitly state when not to use the tool or mention alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_pre_flightC
Read-only
Inspect

Pré-audit comptable — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: Spendesk — Pré-audit commissaire · Readiness 74/100 · 4 findings critiques · Checklist 18 docs. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
auditYes
companyYes
systemsYes
financialsYes
knownIssuesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description adds minimal value beyond confirming a read operation. It mentions server-side validation but lacks detail on side effects or data access patterns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise but includes jargon ('Gapup agent-payable C-suite expertise') that reduces clarity. It is front-loaded but could be more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description does not cover return format or interpretation of the deliverable, and there is no output schema. Given the tool's complexity (6 params, nested objects), more detail is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (17%), and the description only vaguely says 'send the documented case fields' without explaining any parameter semantics. The schema itself provides limited descriptions for nested properties.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs a pre-audit readiness assessment and returns a structured deliverable, referencing a concrete example. However, it does not distinguish this tool from sibling audit-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like esg_audit_multi or privacy_compliance_audit. The reference case provides an example but not contextual decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

banking_fee_negotiatorA
Read-onlyIdempotent
Inspect

As a CFO-focused tool, banking_fee_negotiator analyzes your bank's fee structures (account maintenance, wire transfers, credit lines) and provides data-driven negotiation recommendations. Input your current fees and bank details to receive benchmark comparisons from World Bank and ECB SDW, along with specific levers to reduce costs. Ideal for optimizing treasury operations and improving financial efficiency. Keywords: bank fees, cost optimization, treasury management, financial benchmarking, negotiation strategy.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
industryNoIndustry classification (e.g., 'manufacturing', 'retail')
bank_countryYesISO 2-letter country code of the bank
credit_line_feeNoCurrent annual credit line fee percentage
wire_transfer_feeNoCurrent domestic wire transfer fee in USD
international_wire_feeNoCurrent international wire transfer fee in USD
account_maintenance_feeYesCurrent monthly account maintenance fee in USD

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
negotiation_leversNo
credit_line_benchmarkNoIndustry benchmark for credit line fees percentage
wire_transfer_benchmarkNoRegional benchmark for domestic wire transfer fees in USD
international_wire_benchmarkNoRegional benchmark for international wire transfer fees in USD
account_maintenance_benchmarkNoRegional benchmark for account maintenance fees in USD
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, openWorldHint, and idempotentHint. The description adds that the tool uses World Bank and ECB SDW data for benchmarks and provides specific cost-reduction levers, which is consistent and adds valuable behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph, front-loaded with the core purpose. It is efficient but could be slightly more concise by removing redundant marketing language. Overall, it is well-structured and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (not shown but indicated), the description need not detail return values. It covers the tool's purpose, input requirements, data sources, and use case. For a read-only analytical tool, this is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the input schema already fully describes each parameter. The description repeats some parameter names (account maintenance, wire transfers, credit lines) but does not add significant new meaning beyond what the schema provides. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is a CFO-focused analyzer for bank fee structures, providing negotiation recommendations. It uses specific verbs ('analyzes', 'provides') and distinguishes itself from siblings by being narrowly focused on bank fees and treasury optimization.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (input current fees and bank details to get benchmark comparisons and negotiation levers). It doesn't explicitly state when not to use it, but the niche scope implies it's for bank fee negotiation, which is clear given no closely related sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

battle_cards_liveC
Read-only
Inspect

Fiche de combat live — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub vs McKinsey Lilli — Deal SaaS B2B €500k · Win rate +11 pts · 6 objections clés armées. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
ourOfferYes
competitorYes
dealContextYes
knownWeaknessesNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint and openWorldHint, but the description adds that inputs are validated server-side and output is audited. However, it doesn't disclose any behavioral traits beyond what annotations imply, such as rate limits or side effects. The bar is low due to annotations, but the description adds marginal value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but mixes French and English, includes an unnecessary specific reference case, and uses jargon ('Gapup agent-payable C-suite expertise'). It could be more concise and standardized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters with nested objects and no output schema, the description is insufficient. It doesn't explain what the deliverable contains, how to interpret results, or prerequisites beyond vague reference to 'documented case fields'. The reference case is specific but not generally instructive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'async' has a description). The description does not elaborate on other parameters like competitor, dealContext, or ourOffer. It says 'send the documented case fields' but doesn't map to parameters. With low coverage, description should compensate, but it doesn't.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool returns a structured, audited deliverable referred to as a 'Fiche de combat live' (live battle card). It implies a competitive intelligence output but doesn't explicitly differentiate from siblings like competitive_deep_dive or competitor_intel, limiting clarity in a crowded toolset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The instruction to 'send the documented case fields' is vague and doesn't help an agent decide between this and other competitive analysis tools. Missing when-not or alternative references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

battle_planC
Read-only
Inspect

Plan de bataille marketing — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Gapup Hub — Q3 2026 · Budget €120k · Pipeline €800k · 5 chantiers prioritaires. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
quarterYes
teamSizeYes
arrTargetYes
budgetEurYes
arrCurrentYes
companyNameYes
topChannelsYes
icpDescriptionYes
currentBlockersYes
primaryObjectiveYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and openWorldHint=true, which align with the tool generating a deliverable without side effects. The description adds that inputs are validated server-side, but does not explain the async parameter behavior or the nature of the return value beyond being 'structured, audited'. The added context is minimal and does not disclose potential behaviors like processing time or polling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short (three sentences) but includes a fragmented first sentence and a reference case that may not be universally relevant. It could be more concise by front-loading the core function (marketing plan generation) and removing the cryptic 'Gapup agent-payable C-suite expertise (CMO)' phrase.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 11 parameters, no output schema, and minimal description. The description does not explain the deliverable's structure, the role of the 'async' parameter, or how to interpret results. This leaves significant gaps for the agent to infer or fail on execution.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 9% schema description coverage, the description should compensate by explaining key parameters like companyName, arrCurrent, etc. Instead, it only provides a reference case example, which implies parameter meanings but does not explicitly define them. The schema itself lacks descriptions for most fields, so the tool definition fails to guide the agent on what values to provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool returns a structured, audited marketing deliverable, referencing a 'Plan de bataille marketing' for CMO expertise. However, it does not differentiate from marketing-related siblings like brand_builder or positioning_strategist, leaving ambiguity about when to choose this tool over others.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description lacks explicit use cases, prerequisites, or exclusions, leaving the agent without decision-making support.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bias_amplification_trackerA
Read-onlyIdempotent
Inspect

Tracks bias amplification in LLM outputs by analyzing fairness metrics from HuggingFace's model leaderboard. Designed for risk assessment personas to detect and quantify demographic, gender, or racial bias amplification in generated text. Accepts model identifiers or output samples, returns structured bias metrics and amplification trends.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
modelIdNoHuggingFace model identifier (e.g., 'facebook/opt-1.3b')
outputSamplesNoArray of LLM output strings to analyze for bias amplification
demographicGroupsNoSpecific demographic groups to monitor (e.g., ['gender', 'race'])

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
biasMetricsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint, openWorldHint, idempotentHint, so no destructive actions. The description adds that it accepts model identifiers or output samples and returns structured bias metrics and amplification trends, aligning with annotations and providing useful context beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no fluff: first sentence defines purpose and data source, second sentence specifies users and outputs. Every part is necessary and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's 4 parameters and presence of an output schema, the description covers purpose, inputs, outputs, and target audience. It lacks detail on specific metrics computed, but the output schema likely provides that, so it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains all 4 parameters well. The description adds context by summarizing accepted inputs (modelId, outputSamples, demographicGroups) but does not significantly enhance understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool tracks bias amplification in LLM outputs using HuggingFace model leaderboard, with a specific verb and resource. It distinguishes itself from sibling tools like model_safety_certification_checker or hallucination_confidence_meter by focusing on fairness metrics and demographic bias.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description targets 'risk assessment personas' but does not explicitly state when to use this tool over alternatives or provide exclusions. Usage is implied by its focus on bias amplification, but no direct comparison with sibling tools is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bond_covenant_esg_compliance_checkerA
Read-onlyIdempotent
Inspect

As a CFO, quickly assess whether your bond covenants meet ESG compliance standards set by BIS and ECB. This tool analyzes covenant text against regulatory benchmarks, identifying potential ESG-related risks in carbon emissions, governance practices, and social impact clauses. Input bond covenant details and receive structured compliance insights with source references. Ideal for pre-issuance due diligence or ongoing monitoring of existing bond portfolios.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
couponTypeNoType of bond coupon
covenantTextYesFull text of the bond covenant to analyze
issuerSectorNoIndustry sector of the bond issuer (e.g., energy, finance)
jurisdictionNoLegal jurisdiction governing the bond (e.g., EU, US)
maturityDateNoMaturity date of the bond

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
riskAreasNo
complianceScoreNo
recommendationsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, indicating safe, idempotent behavior. The description adds value by specifying that the output includes 'structured compliance insights with source references,' providing additional clarity on what the agent can expect beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long, front-loaded with the primary purpose, and contains no filler. Every sentence adds value: context (CFO role), mechanism (analyzes against benchmarks), and use cases (pre-issuance/monitoring). It is concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's purpose, usage context, and output nature. Given that the input schema and output schema exist, the description provides sufficient high-level context. It omits details on specific output fields, but the output schema handles that, so completeness is solid.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the schema already documents all 6 parameters adequately. The description does not add significant extra meaning about the parameters beyond what is in the schema, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to assess bond covenant compliance against BIS and ECB ESG standards. It specifies the verb 'assess' and the resource 'bond covenant text' with regulatory benchmarks. However, it does not explicitly distinguish itself from similar sibling tools like bond_covenant_monitor or esg_audit_multi, preventing a score of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage context, stating it is 'Ideal for pre-issuance due diligence or ongoing monitoring.' This gives general guidance but lacks explicit when-to-use/when-not-to-use instructions or comparisons to alternative tools, earning a score of 3.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bond_covenant_monitorB
Read-onlyIdempotent
Inspect

As a CFO, monitor bond covenant compliance by analyzing leverage ratios (debt-to-equity, debt-to-EBITDA) and interest coverage ratios using real-time financial data. Input a company's ticker symbol and optional covenant thresholds to receive compliance status, key financial metrics, and SEC filing references. Ideal for proactive debt management and regulatory compliance tracking. Keywords: bond covenants, leverage ratio, interest coverage, debt compliance, SEC filings, financial health.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
tickerYesCompany ticker symbol (e.g., 'AAPL')
covenantThresholdsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
warningsYes
debtToEquityNo
leverageRatioNo
lastFilingDateNo
complianceStatusYes
interestCoverageNo
nextFilingDeadlineNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds context about using real-time data and SEC references but does not contradict annotations. It provides some behavioral context beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences plus a keyword list. It is moderately concise but contains marketing fluff ('As a CFO', 'Ideal for proactive debt management') that could be trimmed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description appropriately notes the returned items (compliance status, metrics, SEC refs). It covers input, output, and use case, leaving minimal gaps for its complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (67% with descriptions for all parameters and nested property descriptions). The description mentions 'debt-to-EBITDA' which is not a parameter, adding slight confusion. It does not significantly add meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool monitors bond covenant compliance using leverage and interest coverage ratios. However, it does not explicitly distinguish itself from the sibling 'bond_covenant_esg_compliance_checker', which could cause confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions input requirements (ticker, optional thresholds) and use case (debt management, compliance tracking), but lacks explicit guidance on when not to use this tool or mention of alternatives like the ESG-focused sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bp_narratifC
Read-only
Inspect

Business Plan narratif — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: Stripe Series A 2012. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
raiseYes
companyYes
keyMetricsYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint and openWorldHint, but the description adds minimal behavioral context. It does not describe output format, size limits, or side effects beyond stating it returns a deliverable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose and reference case. Efficient but could add more useful detail without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 parameters, nested objects, no output schema, and many siblings, the description lacks explanation of output, async behavior, and differentiation from similar tools like ftg_business_plan.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (25%) with only async described. The description says inputs are validated server-side but provides no additional meaning for the complex nested parameters (company, raise, keyMetrics).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns a structured, audited business plan narrative with CFO expertise, and gives a reference case. However, it does not differentiate from the sibling tool ftg_business_plan, which may have similar purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. The description only mentions server-side validation and to 'send the documented case fields', without context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brand_builderC
Read-only
Inspect

Architecte de marque — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Pennylane — brand identity SaaS fintech B2B FR/EU (2023). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
brandYes
targetYes
founderYes
existingAssetsNo
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds only that inputs are validated server-side and returns a deliverable, but does not detail behavioral aspects like auth needs, rate limits, or processing time.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description mixes French and English, includes a specific reference case that is not immediately helpful, and is not front-loaded with the most critical information. The space is not fully justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and complex nested parameters, the description should clarify the return format and expected output. It merely says 'returns a structured, audited deliverable' without specifics, leaving the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (20%). The description does not explain the parameters beyond 'send the documented case fields', failing to add meaning for the complex nested objects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a brand architect tool providing C-suite expertise and returns a structured deliverable. It references a specific case (Pennylane) but does not explicitly differentiate from siblings like brand_equity_voice_share_calculator or positioning_strategist.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use vs. alternatives or when not to use. The description mentions 'send the documented case fields' but lacks exclusions or context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brand_equity_voice_share_calculatorA
Read-onlyIdempotent
Inspect

Calculates brand equity voice share for CMOs by analyzing mentions across 500K+ news articles and forums from Common Crawl and Wayback Machine. Inputs include brand name, competitors, and time range. Outputs voice share percentage, sentiment distribution, and top sources. Ideal for competitive benchmarking and brand visibility tracking. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
brandYes
time_rangeYes
competitorsNo
include_forumsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
top_sourcesNo
total_mentionsNo
brand_voice_shareNo
sentiment_distributionNo
competitor_voice_sharesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, open world, and idempotent behavior. The description adds context about data sources (Common Crawl, Wayback Machine) and potential timeouts, and outlines outputs (voice share percentage, sentiment distribution, top sources). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: three short sentences that front-load the primary purpose and quickly cover key inputs, outputs, and a usage tip. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters (2 required), nested objects, and an output schema, the description provides a high-level overview but lacks detail on parameter semantics and behavior for optional parameters. The existence of an output schema helps, but the description does not fully equip an agent to correctly populate all inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (20%), so the description should compensate. While it mentions brand, competitors, time range, and async flag, it fails to explain the 'competitors' array, 'include_forums' boolean, or the structure of the 'time_range' nested object. This leaves significant gaps for an agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool calculates brand equity voice share for CMOs by analyzing mentions across 500K+ news articles and forums. It specifies inputs, outputs, and data sources, distinguishing it from sibling tools like 'brand_builder' or 'content_engine'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides guidance on when to use the tool ('Ideal for competitive benchmarking and brand visibility tracking') and hints at timeout behavior ('Pass async:true to avoid timeout'). However, it does not explicitly state when not to use or mention alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

budget_variance_aiB
Read-only
Inspect

Analyse d'écart budgétaire — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Answers: Explain the key drivers of the budget vs actual variance for in — what are the top 10 narrative explanations? · Which cost categories drove the budget overrun for in , and what corrective actions should management take? · Revise the Q4 forecast based on observed Q3 variances for — give me 3 scenarios (base, optimistic, conservative). · Prepare a board-ready budget variance memo for , budget €M vs actual €M, with management actions. · What are the quick wins to reduce budget overspend for by end of quarter without impacting growth targets? Reference case: Doctolib Q3 2026 — budget €38.5M vs actual €41.2M (+7.0%) — cloud + headcount + deals timing. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
entityYes
budgetContextYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, openWorldHint), the description adds that inputs are validated server-side and that it returns a structured, audited deliverable. This provides useful behavioral context without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately concise but includes a verbose list of example queries and mixes French and English. It could be more structured by separating core function from examples.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 parameters, nested objects, no output schema) and many siblings, the description lacks details on return format, error handling, async behavior documentation (only in schema), and how to interpret results. It feels incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only async parameter documented). The description does not explain the meaning of entity or budgetContext parameters, despite their nested structure and required fields. Users must infer from the example queries.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a budget variance analysis tool that returns a structured deliverable, with example queries specifying its capabilities. However, it does not sharply distinguish it from similar finance tools like earnings_reviewer or financial_model_3statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description provides example queries but lacks context about prerequisites, limitations, or comparisons with sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

building_enrichAInspect

Enrich a location with European building intelligence: roof surfaces (m²), parking areas, solar-obligation status under French loi APER and loi Climat-Résilience, existing solar installations. Covers 628,000 scanned roofs and 91,800 parkings across 6 EU countries (FR, DE, IT, ES, BE, NL). Deterministic database lookup — no LLM, no generation, sub-second.

ParametersJSON Schema
NameRequiredDescriptionDefault
latYesLatitude (WGS84)
lngYesLongitude (WGS84)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
radiusMNoSearch radius in metres (default 150, max 500)
minAreaM2NoOnly return roofs/parkings at least this large, in m²

Output Schema

ParametersJSON Schema
NameRequiredDescription
queryNo
roofsNo
summaryNo
parkingsNo
solarInstalledNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states 'Deterministic database lookup — no LLM, no generation, sub-second,' which conveys predictable, fast, non-generative behavior. It also provides coverage statistics, giving a sense of limitations. However, it does not describe behavior for unmatched locations or potential error modes, which would make it fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured, with the main purpose stated in the first sentence, followed by coverage data and behavioral characteristics. Every sentence adds value: purpose, scope/coverage, and performance/reliability. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having 5 parameters and no annotations, the description is sufficiently complete for a database lookup tool. It covers purpose, geographic scope (6 EU countries), specific data categories, legal context, and performance characteristics. The existence of an output schema means return values need not be described in text. The only minor gap is explicit usage alternatives, but that is already covered under usage guidelines.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all 5 parameters, including lat/lng, async, radiusM, and minAreaM2. The tool description adds context like 'roof surfaces (m²)' and 'parking areas' that aligns with minAreaM2, but it does not substantially exceed what the schema already documents. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a specific action verb ('Enrich') and identifies a clear resource ('a location with European building intelligence'), then lists concrete data types (roof surfaces, parking areas, solar-obligation status, existing solar installations). This clearly distinguishes it from broader sibling tools like real_estate_intel or geo_logistics_intel by specifying the exact domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context through mentions of French solar obligations and EU coverage, but does not explicitly state when to use this tool versus alternatives or when not to use it. There are no exclusions or alternative tool references, leaving usage guidance implicit rather than direct.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

candidate_screening_rankingA
Read-onlyIdempotent
Inspect

AI-powered candidate screening and ranking for recruiters, hiring managers, ATS providers and recruitment AI agents. Ingests a job description and 1-50 candidate resumes, returning a ranked shortlist with score breakdowns across five weighted criteria: skills_match (tech stack and soft skills extracted from JD vs resume), experience_match (years vs seniority level inferred from JD), education_match (degree level + top-school detection), role_progression (Junior to Senior to Lead patterns), culture_fit_estimate (remote/hybrid, startup vs enterprise). Per candidate: overall_score 0-100, matched/missing skills, red_flags (job hopping, employment gaps, seniority mismatch), green_flags (long tenure, promotions), 3-5 interview questions, fit_summary. Diversity signals are first-name proxies ONLY with mandatory ethical WARNING. All processing is local -- no external API calls, instant response, privacy-preserving.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
candidatesYesArray of candidate objects. Maximum 50.
role_countryNoOptional ISO 2-letter country code for regional context (informational).
job_descriptionYesFull text or summary of the job description and role requirements.
criteria_weightsNoOptional weighting per criterion. Default: skills=0.4, experience=0.2, education=0.1, progression=0.15, culture=0.15.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
nice_to_haveYes
quality_scoreYes
required_skillsYes
candidates_rankedYes
diversity_signalsNo
shortlist_recommendedYes
job_description_summaryYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond annotations by explicitly stating that all processing is local, privacy-preserving, instant, and that diversity signals are first-name proxies only with a mandatory ethical warning. It also describes the scoring breakdown per criterion and the output features (red flags, green flags, interview questions). This adds significant behavioral context that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense yet well-structured. It front-loads the purpose, then lists criteria, output format, and caveats. Every sentence adds value without redundancy, achieving conciseness despite the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, nested objects, output schema), the description covers all necessary aspects: input requirements, processing details, scoring criteria, output fields (including ethical warnings), and performance characteristics. It leaves no obvious gaps for an agent to misunderstand usage or expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the baseline is 3, but the description adds substantial meaning beyond the schema by explaining the five weighted criteria in detail (e.g., skills_match extracts tech stack and soft skills, experience_match compares years to seniority level). This extra context helps the agent understand how to effectively set criteria_weights and interpret results.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: AI-powered candidate screening and ranking that ingests a job description and 1-50 resumes to return a ranked shortlist with score breakdowns across five weighted criteria. It uses specific verbs and resources ('screening and ranking', 'ingests', 'returning') and provides detailed criteria, effectively distinguishing it from sibling tools like 'talent_intelligence' or 'recruiting_architect'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly defines usage by detailing the tool's input and output, making it clear that it is intended for screening and ranking candidates. However, it does not explicitly state when to use this tool versus alternatives, nor does it mention exclusion criteria or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capacity_planningC
Read-only
Inspect

Planification capacitaire — Gapup agent-payable C-suite expertise (CHRO). Returns a structured, audited deliverable. Reference case: Gapup Hub — 22→48 FTE en 12m · ARR €480k→€1.7M · Plan d'embauches par département. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
benchmarksNo
financialsYes
constraintsNo
currentTeamYes
hiringBudgetEurNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side, implying error handling behavior. It does not contradict annotations, but it adds minimal behavioral context beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief (3 sentences) but mixes French and English, which may reduce clarity for some agents. It front-loads the purpose and includes a reference case, but the language mixing is a minor structural flaw.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema, complex nested inputs, and 7 parameters, the description is insufficient. It doesn't describe the deliverable's structure, how to interpret results, or any expected response format. The reference case hints at outcomes but not structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14%, and the description contributes no parameter-level meaning. It vaguely says 'send the documented case fields' but doesn't list or explain parameters. The schema's own descriptions for nested objects are sparse, leaving the agent with little guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly indicates the tool is for capacity planning for CHROs, mentioning it returns a structured, audited deliverable and providing a reference case. The French title and description reinforce the specific domain, though it lacks a single verb+resource phrase. The purpose is distinct from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives like 'capital_strategy' or 'growth_path_architect'. It only asserts inputs are validated server-side, which is a technical note, not usage context. No exclusions or examples of when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capital_strategyC
Read-only
Inspect

Stratégie de financement — Gapup agent-payable C-suite expertise (CSO). Returns a structured, audited deliverable. Reference case: Alan assurance santé SaaS — séquence Seed→A→B→C (2016-2022). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
growthPlanYes
financialPositionYes
founderConstraintsYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds little beyond stating it returns a structured deliverable and that inputs are validated server-side. No disclosure of side effects or constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise at 4 sentences, front-loaded with the purpose. However, it could be more informative without adding length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested objects, 5 parameters, no output schema), the description is incomplete. It does not explain the deliverable's contents, how to interpret results, or mention the async parameter. The reference case provides some context but insufficient detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, and the tool description does not describe any parameters. It only says 'send the documented case fields' without detailing the required fields, leaving the agent without sufficient context to populate inputs correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool returns a structured, audited deliverable for financing strategy, and provides a reference case. However, it does not explicitly differentiate from sibling tools, though the purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lacks explicit guidance on when to use this tool versus alternatives. It does not mention when not to use it or provide comparison with siblings, only stating that inputs are validated server-side.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cap_table_strategistC
Read-only
Inspect

Stratège du cap table — Gapup agent-payable C-suite expertise (FUNDRAISING). Returns a structured, audited deliverable. Reference case: Aleph AI Series B — modèle dilution multi-rounds + simulations secondaires + hygiène equity · 5 scenarios. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
plannedRoundsYes
currentCapTableYes
founderObjectivesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the description adds minimal behavioral context beyond stating it returns a deliverable. No contradictions, but no rich elaboration on processes or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences in French plus a reference case. It is somewhat concise but could be more structured, and the mix of French and English context may hinder clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested objects, no output schema), the description is insufficient. It does not explain the deliverable's content, interpretation, or prerequisites, leaving many gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 17% schema description coverage, the description should compensate by explaining key parameters, but it does not. It only says 'send the documented case fields', providing no additional meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is a cap table strategist for fundraising, returning a structured audited deliverable. It mentions a reference case, which aids understanding, but does not explicitly differentiate it from sibling tools like 'capital_strategy' or 'deal_structurer'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description lacks context about appropriate scenarios, exclusions, or comparisons to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

carbon_footprint_calculatorB
Read-onlyIdempotent
Inspect

Calculate a company's greenhouse-gas footprint under the GHG Protocol (Scope 1 + 2 + 3, in tCO2eq, tier-2 accuracy ±20%). Returns the emissions breakdown, hotspot identification, 5-8 reduction levers each with capex and payback, an SBTi-aligned reduction trajectory over 5-25 years, the 15 Scope-3 categories in detail, and CSRD/ESRS reporting readiness. When to use this tool: the user needs a carbon assessment for CSRD compliance pre-audit, green-finance access, or supplier ESG scorecards. Inputs: the company profile and its activity data. Delivered by Émilie, the AI Sustainability lead of the Gapup portfolio.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
perimeterYes
scope1SourcesNo
scope2SourcesYes
reductionTargetsNo
scope3ActivitiesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
kpisNo3-5 headline ESG KPI bubbles
hotspotsYesTop emission sources ranked by contribution
breakdownYesEmissions breakdown by scope
csrdReadinessYesCSRD/ESRS reporting readiness assessment
sbtiTrajectoryNoSBTi-aligned annual reduction trajectory
reductionLeversYes5-8 actionable reduction levers with financial analysis
executiveSummaryYesBoard-ready GHG assessment prose
scope3CategoriesNoGHG Protocol 15 Scope-3 categories detail
totalEmissionsTco2eqYesTotal GHG footprint in tCO2eq (Scope 1+2+3 combined, ±20% tier-2 accuracy)
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, idempotentHint, destructiveHint) already cover safety and idempotency. The description adds accuracy (±20%) and output details but does not disclose behavioral traits like execution time or side effects beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single coherent paragraph with logical flow: what it does, outputs, when to use, inputs. However, the 'Delivered by Émilie' line is extraneous and adds no value, preventing a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 8 parameters and low schema coverage, the description provides a good overview of outputs and purpose but lacks detail on input parameters and their relationships. The existence of an output schema helps, but gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13%, and the description only vaguely mentions 'company profile and its activity data.' It fails to elaborate on the eight parameters, including nested objects like company, perimeter, and emission sources, leaving the agent with insufficient guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool calculates a company's greenhouse-gas footprint under the GHG Protocol, specifying scope, accuracy, and outputs. It distinguishes from siblings like 'esg_audit_multi' by focusing on carbon footprint with precise protocol and tiers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides use cases: 'CSRD compliance pre-audit, green-finance access, or supplier ESG scorecards.' It offers clear context but does not specify when not to use or mention alternative tools, which slightly reduces the score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

carbon_roadmapC
Read-only
Inspect

Roadmap carbone — Gapup agent-payable C-suite expertise (SUSTAINABILITY). Returns a structured, audited deliverable. Reference case: Cas démo — Roadmap carbone. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
perimeterYes
scope1SourcesNo
scope2SourcesYes
reductionTargetsNo
scope3ActivitiesNo
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true. The description adds no additional behavioral context, such as authentication requirements, rate limits, or what happens to existing data. It merely restates that it returns a deliverable, which is already implied by the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (two sentences) but includes unnecessary jargon ('Gapup agent-payable C-suite expertise') that does not aid understanding. The second sentence about validation and reference case is useful but could be more concise. Overall adequate but not efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, nested objects, no output schema), the description is severely lacking. It does not mention return format, output structure, or any examples. The absence of output schema increases the need for description completeness, which is not met.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13%, yet the description provides zero explanation of parameters or their meaning. It only says 'send the documented case fields,' which adds no value. The parameters are complex nested objects, and without description assistance, agents cannot understand usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it 'returns a structured, audited deliverable' related to a carbon roadmap, but lacks a clear verb-resource pairing (e.g., 'Generates a carbon reduction roadmap'). The jargon 'Gapup agent-payable C-suite expertise' obscures the purpose. Among sibling tools like 'carbon_footprint_calculator' and 'sustainability_report', this description does not differentiate clearly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description mentions input validation and a reference case but does not specify when it is appropriate or not appropriate to invoke this tool compared to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

champion_mappingC
Read-only
Inspect

Cartographie du champion — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Spendesk × Decathlon (deal €120k/an) — Champion identifié : CFO Group · Plan 6 semaines multi-touch. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
dealYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
knownContactsYes
sellerContextYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint and openWorldHint. Description adds that inputs are validated server-side and output is a structured audited deliverable, which is consistent. No additional behavioral traits disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, but the reference case adds length without essential information. Could be trimmed to be more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and complex nested parameters; the description does not explain the output format or the purpose of individual fields, leaving gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, and the description does not explain the parameters beyond 'send the documented case fields'. It fails to add semantic value for the 4 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it performs champion mapping for C-suite expertise and returns a structured deliverable. The reference case helps illustrate, though the exact scope is slightly vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternative tools. Sibling tools suggest many analysis options, but no differentiation is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

change_failure_root_cause_classifierA
Read-onlyIdempotent
Inspect

Classifies root causes of change failures for CTO-level incident analysis. Uses GitHub PR metadata and Snyk vulnerability data to identify patterns like dependency vulnerabilities, configuration drift, or deployment process gaps. Inputs include GitHub PR URL or incident ID, and outputs structured root cause categories with confidence scores. Ideal for post-mortem analysis and change risk assessment.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
pr_urlYes
incident_idNo
snyk_org_idNo
time_range_daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
root_causesNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint, openWorldHint, and idempotentHint, indicating safe, non-destructive behavior. The description adds that it uses specific data sources and produces categories with confidence, but does not disclose additional behavioral traits like failure modes or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences with no redundancies. The description is well-structured and front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main purpose and general inputs but omits details on optional parameters like snyk_org_id and time_range_days. Given the tool has 5 parameters and an output schema, more parameter context would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (20%), and the description does not mention any parameter details, such as required fields or valid values. It fails to compensate for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool classifies root causes of change failures using GitHub PR metadata and Snyk vulnerability data. It specifies the output as structured categories with confidence scores, making the purpose distinct from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says 'Ideal for post-mortem analysis and change risk assessment,' implying usage context but not explicitly stating when to use or alternatives. No exclusion criteria are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

china_ecommerce_intelA
Read-only
Inspect

Chinese e-commerce intelligence for the ZH diaspora (50M+), import-export teams, brand IP enforcement, MENA/Africa entrepreneurs sourcing from China, and brand monitoring. Covers Taobao, Tmall, JD.com, Pinduoduo, 1688.com (B2B) and AliExpress (cross-border).

Five modes: • product_search — search products by keyword across CN platforms. Returns title ZH/EN, price CNY + USD estimate, sales 30d, rating, seller info, product URL. • seller_profile — full seller/supplier dossier: factory vs reseller detection, certifications (ISO, BSCI, CE), rating, years in business, main categories. • price_history — 12-month price trend for a product (live current price + seasonal model for CN shopping festivals: 11.11, 6.18, CNY). • brand_monitoring — detect counterfeits and grey market listings: price anomaly detection (>50% below MSRP = suspicious), counterfeit keyword scan, risk score 0-100. • market_intel — category overview: top 5 sellers by market share, avg/median price, volume estimate, price range.

Data quality note: LIVE data from Taobao/Tmall/JD/Pinduoduo REQUIRES AICI_RESEARCH_PROXY_URL with CN residential routing (Bright Data -country-cn). Without proxy: AliExpress (cross-border) + curated category fallback available.

Input formats for seller_profile: 'platform:id' e.g. 'aliexpress:123456', '1688:87654321', 'tmall:apple-store-official'. Input formats for price_history: AliExpress product URL or numeric product ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesAnalysis mode. product_search=find products, seller_profile=supplier dossier, price_history=price trend, brand_monitoring=counterfeit detection, market_intel=category overview.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesKeyword, product name, product_id, seller_id (platform:id), brand name, or category. Accepts Chinese characters (ZH) or English.
regionNoMarket region. CN-domestic=full platform coverage, cross-border=AliExpress+1688 focus. Default: CN-domestic.
platformNoTarget platform. Default: all. Note: taobao/tmall/jd/pinduoduo require CN proxy.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
signalsYes
sourcesYes
productsNo
market_intelNo
platform_usedYes
price_historyNo
quality_scoreYes
seller_profileNo
brand_monitoringNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint, openWorldHint. Description adds valuable behavioral context: live data requires CN proxy, otherwise fallback to AliExpress; async mode described. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured with bullet points for modes and notes. It's slightly long but every sentence adds value. Could be more concise, but not overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given high schema coverage, presence of output schema, and annotations, description covers modes, inputs, data sources, and limitations. Missing error handling details, but satisfactory overall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3). Description adds meaning beyond schema: input formats for seller_profile and price_history, region and platform notes, and query maxLength. Enhances clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Chinese e-commerce intelligence for the ZH diaspora...' and lists five specific modes with brief explanations. It distinguishes from siblings like 'china_market_data' by being more comprehensive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides guidance on when to use each mode, data quality notes (proxy requirement), and input formats. However, it lacks explicit when-not-to-use instructions or alternatives for similar tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

china_market_dataA
Read-only
Inspect

Chinese capital market intelligence for the ZH diaspora (50M+) and institutional investors. Covers A-Shares (SSE/SZSE), H-Shares (HKEX), and ADRs across four modes:

• company — full company profile: name ZH/EN, USCC (18-digit social credit code), exchange, industry (CSRC classification), chairperson, registered capital, SOE flag • market_quote — real-time quote: price (CNY or HKD), change%, volume, market cap, P/E ratio, dividend yield, last update timestamp • sector_overview — sector snapshot: top 5 companies by market cap, avg P/E, 30-day sector index change. Supported sectors: semiconductor, ev, battery, technology, finance, energy, realestate, consumer, pharma, telecom • regulatory_filing — recent regulatory disclosures (HKEX filings: annual, quarterly, announcements, mergers, IPOs) with title, date, document URL

Input formats accepted: • 6-digit A-Share ticker (e.g. '600519' for Moutai SSE) • HKEX ticker (e.g. '0700.HK' or '700' for Tencent) • Company name in EN or ZH (e.g. '腾讯', 'Kweichow Moutai') • Sector keyword (e.g. 'semiconductor', '半导体')

Data sources: Yahoo Finance (primary, always accessible), Eastmoney push2 + CompanySurvey (via Bright Data proxy when AICI_RESEARCH_PROXY_URL is set), HKEX filing API. Note: Eastmoney/CSRC/SSE are blocked from datacenter IPs without proxy — set AICI_RESEARCH_PROXY_URL to unlock full coverage.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesAnalysis mode. company=full profile, market_quote=price data, sector_overview=top 5 by sector, regulatory_filing=recent filings.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesTicker (6-digit A-share, 4-digit HK, Yahoo format), company name (ZH or EN), or sector keyword.
exchangeNoExchange filter. Default: all. Affects sector_overview ticker selection.
period_daysNoLookback period in days for regulatory filings. Default: 30.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
queryYes
statusYes
companyNo
sourcesYes
market_quoteNo
quality_scoreYes
sector_overviewNo
regulatory_filingsNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral context beyond annotations: it details data sources (Yahoo Finance, Eastmoney, HKEX), proxy requirements for full coverage, and the asynchronous polling mechanism via job_result. No contradiction with annotations (readOnlyHint=true, destructiveHint=false) is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with sections and bullet points, making complex information scannable. It is slightly verbose but every sentence adds clarity given the tool's breadth. No redundancy is present, earning a high score for its informational density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, 2 required, 2 enums, output schema exists), the description covers all essential aspects: modes, input formats, data sources, proxy requirements, and the async flow. It is complete without needing to describe return values, as the output schema handles that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by elaborating on the modes (e.g., 'company — full company profile...'), clarifying input formats (ticker, name, sector), and explaining the async parameter’s purpose. This extra context raises the score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it provides Chinese capital market intelligence, covering A-Shares, H-Shares, and ADRs across four distinct modes (company, market_quote, sector_overview, regulatory_filing). Each mode is briefly defined with specific output details, making the tool's purpose highly specific and distinguishable from siblings like india_market_data or historical_price_series.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (for Chinese market data) but does not explicitly state when not to use it or name alternative tools. It lacks explicit guidance on distinguishing from similar tools (e.g., india_market_data, realtime_data_streams), leaving the agent to infer usage boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

churn_defenderB
Read-only
Inspect

Bouclier anti-churn — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Spendesk — portefeuille 400 clients PME/ETI, détection churn Q2 2025 (€8M ARR). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
accountsYes
csrContextNo
analysisWindowDaysYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, indicating no side effects. The description adds that the tool returns a 'structured, audited deliverable' and mentions server-side validation, which is consistent. However, it does not elaborate on what the deliverable contains, authentication needs, or rate limits. The behavioral insight is adequate but not enriched beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, reasonably concise and front-loaded with the tool's identity. The reference case adds some length but is relevant. It could omit the reference case or integrate it more efficiently, but overall it wastes little space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested objects, 5 parameters, no output schema) and the presence of many related sibling tools, the description is insufficient. It does not explain the deliverable's structure, how results are audited, or how to handle errors. The absence of output schema and the high parameter count demand a richer description to guide the agent effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not explain any parameters beyond the schema; it merely says 'send the documented case fields.' With schema coverage at only 20% (only async described), the description fails to compensate by clarifying the purpose of the many nested fields (e.g., company properties, account signals). A tool with complex inputs requires more parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as an anti-churn analysis tool that returns a structured deliverable for C-suite (CRO). The reference case adds specificity. However, it does not explicitly differentiate from sibling tools like renewal_optimizer, save_plays, or upsell_hunter, which operate in adjacent spaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a concrete reference case (Spendesk) indicating a typical use case, but lacks explicit guidance on when to use this tool versus alternatives, or any exclusion criteria. The phrase 'send the documented case fields' implies a specific input format but does not define the context of use relative to other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

climate_scenario_rcpA
Read-only
Inspect

Projections climatiques long terme par scénario IPCC (RCP AR5 + SSP AR6) pour toute localisation. Scénarios : RCP_4_5, RCP_8_5 (AR5), SSP1_2_6, SSP2_4_5, SSP3_7_0, SSP5_8_5 (AR6), ou 'all' (compare tous). Horizons : 2030–2100. Métriques : température (delta vs baseline 1990-2010, jours >35°C, nuits chaudes), précipitations (delta%, événements extrêmes, sécheresses), hausse du niveau de la mer (cm vs 2000), événements extrêmes (ouragans, inondations P100, sécheresses), indice incendie. Sorties : comparaison multi-scénarios, probabilité IPCC, signaux d'impact business par secteur. Sources : Open-Meteo CMIP6 (keyless), IPCC AR6 Atlas lookup, NOAA SLR projections. Usages : TCFD/CSRD physical risk, due diligence actifs long terme, assurance catastrophe, planification infrastructure. Cache 7j. SLA ≤20s.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
metricsNoMétriques à inclure. Défaut : toutes.
locationYesLocalisation : {city, country?} ou {lat, lon}
scenarioYesScénario IPCC. 'all' génère une comparaison multi-scénarios.
horizon_yearYesAnnée horizon de la projection (2030–2100)
compare_baselineNoComparer vs baseline 1990-2010 (défaut true)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
locationYes
scenarioYes
projectionsYes
horizon_yearYes
quality_scoreYes
baseline_periodNo
ipcc_likelihood_labelYes
business_impact_signalsYes
multi_scenario_comparisonNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations show readOnlyHint=true, destructiveHint=false, idempotentHint=false. The description adds behavioral details: caching (7 days), SLA (≤20s), and async behavior. This is consistent and provides useful operational context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that front-loads the main purpose and uses efficient, information-dense sentences. Could benefit from bullet points for readability, but every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters, nested objects, enums, output schema), the description covers all necessary aspects: scenarios, metrics, horizon, location, use cases, data sources, caching, and SLA. The output schema exists so return value details are not required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so description adds value by explaining parameter context: what 'all' scenario does, metric descriptions (temperature delta vs baseline, extreme events), and location flexibility (city/country or lat/lon). This enriches understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: long-term climate projections under IPCC scenarios (RCP and SSP) for any location. It lists specific scenarios, metrics, horizons, and output types, distinguishing it from sibling tools like weather_climate_intel by focusing on IPCC-based projections for TCFD/CSRD risk assessment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit use cases (TCFD/CSRD physical risk, due diligence, insurance, infrastructure planning) and mentions multi-scenario comparison. However, it does not explicitly exclude or compare to other tools like weather_climate_intel, leaving room for ambiguity about alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clinical_evidence_brieferB
Read-only
Inspect

Brief évidence clinique (GRADE) — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Answers: Review the clinical evidence for <drug/intervention> in — GRADE rating, key trials, safety signals. · Scan safety signals for in — adverse events, severity, frequency from FAERS and trial data. · Assess comparative effectiveness of versus for — what does the evidence show? · Is there evidence supporting drug repurposing of for — existing trials and GRADE quality? · What are the evidence gaps for in before formulary adoption? Reference case: Semaglutide 2.4mg · Chronic weight management in non-diabetic adults · GRADE high efficacy · studies found · nausea/GI signals · FDA approved · PubMed+ClinicalTrials+OpenFDA. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
topicYes
max_studiesYes
interventionNo
evidence_focusYesall
target_diseaseNo
date_range_yearsYes
intervention_typeNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds context: inputs are validated server-side, it returns a structured/audited deliverable, and async usage is explained. No contradictions; the extra detail on validation and async behavior is helpful beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose but becomes verbose with multiple example questions and a reference case. It could be more concise by focusing on the tool's role and key constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema is provided, and the description does not detail the return value structure beyond 'structured, audited deliverable'. With 8 parameters and 4 required, the description omits critical info like output fields or how results are formatted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (13%). The description does not explain key parameters like topic, max_studies, evidence_focus, or intervention_type. It only vaguely mentions 'send the documented case fields'. With 8 parameters and minimal schema descriptions, the description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it produces a clinical evidence briefing with GRADE ratings, and gives specific example queries (e.g., 'Review the clinical evidence for <drug/intervention> in <indication>'). It distinguishes from siblings like clinical_pharma_intel by focusing on GRADE and structured deliverables. However, jargon ('Gapup agent-payable C-suite expertise') slightly muddies clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage via example questions (e.g., safety scanning, comparative effectiveness) but does not explicitly state when to use this tool versus alternatives like sci_literature_search or clinical_pharma_intel. No when-not-to-use guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clinical_pharma_intelA
Read-only
Inspect

Clinical and pharmaceutical intelligence for biotech analysts, healthcare fund managers, pharma BD teams, catalyst-driven hedge funds and health journalists. Aggregates live data across five modes: • trials — active/completed clinical trials (ClinicalTrials.gov v2 + EU CTR in parallel, 450k+ records) • pipeline — full pipeline by sponsor: trial count by phase + top indications • approvals — FDA drug label approvals + mechanism of action (OpenFDA) • recalls — FDA enforcement recalls classified by severity (Class I/II/III) • adverse_events — FAERS aggregated reactions: top 10 reactions + serious%

Signal detection (P0/P1/P2): P0 if Class I recall OR trial terminated for safety reason P1 if serious adverse events >30% OR ≥3 recalls in 12 months P2 otherwise (standard monitoring)

All sources are public and keyless. Optional env OPENFDA_API_KEY raises daily quota from 1,000 to 120,000 requests. SLA: ≤16s p95 (parallel fetch, 8s budget per source). Cache: 6h trials, 24h approvals, 12h recalls, 6h adverse events.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoAnalysis mode. Default "trials". trials=clinical trials, pipeline=sponsor overview, approvals=FDA approvals, recalls=enforcement, adverse_events=FAERS
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
phaseNoFilter trials by phase (1/2/3/4/NA). Only applies to modes trials and pipeline.
queryYesDrug name, indication, sponsor or molecule (e.g. "atezolizumab", "metastatic NSCLC", "Roche", "semaglutide")
countryNoISO 2-letter country code to filter trial sites (e.g. US, FR, DE).
max_resultsNoMaximum number of results to return. Default 20.
status_filterNoFilter trials by status. Only applies to modes trials and pipeline.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
queryYes
statusYes
trialsNo
recallsNo
signalsYes
sourcesYes
pipelineNo
approvalsNo
quality_scoreYes
adverse_eventsNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnlyHint=true, destructiveHint=false), the description adds significant behavioral context: public keyless sources, optional API key for quota, SLA (≤16s p95), cache durations per mode, and signal detection schema (P0/P1/P2). This fully discloses operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with sections for modes, signal detection, and technical details. It front-loads the purpose and target users. While somewhat lengthy, every sentence adds value and the structure aids readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 params, 5 modes, output schema), the description covers all essential aspects: mode functions, signal detection, source info, caching, SLA, and optional API key. It leaves no significant gaps for an AI agent to correctly invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description repeats some parameter info (e.g., mode options, default) but adds context like 'trials' default mode and signal detection. No major additional semantics beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides clinical and pharmaceutical intelligence, listing five distinct modes (trials, pipeline, approvals, recalls, adverse_events) with specific sources and purposes. This differentiates it from other tools in the sibling list, which cover unrelated domains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly signals usage through target audience (biotech analysts, healthcare fund managers, etc.) and mode definitions. However, it does not explicitly state when to use this tool versus alternatives or provide exclusion cases, leaving some gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cloud_cost_ri_optimizerA
Read-onlyIdempotent
Inspect

Analyzes AWS and Azure cloud pricing data alongside RIPE regional demand trends to generate Reserved Instance purchase recommendations for CTOs. Inputs include target cloud provider, instance family, region, and desired commitment term. Outputs include cost savings percentage, optimal RI quantity, and regional demand insights. Ideal for reducing cloud spend with data-driven decisions. Keywords: cloud cost optimization, reserved instances, AWS pricing, Azure pricing, RIPE demand trends.

ParametersJSON Schema
NameRequiredDescriptionDefault
termNo
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
regionYes
utilizationNo
cloud_providerYes
instance_familyYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
ri_costNo
sourcesNo
warningsNo
on_demand_costNo
break_even_monthsNo
regional_demand_scoreNo
cost_savings_percentageNo
recommended_ri_quantityNo
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate read-only, idempotent, and open-world, so the safety profile is clear. However, the description fails to mention the 'async' parameter defined in the schema, which is a key behavioral aspect for handling slow queries. This omission reduces transparency significantly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (four sentences plus keywords) and front-loaded with the main action. It uses a clear list structure for inputs and outputs. Slight redundancy from the keywords section, but no significant waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given six parameters and an output schema exists, the description covers the essential inputs and high-level outputs. It lacks details on async usage, prerequisites, data freshness, or error handling. Adequate but not comprehensive; the output schema helps but does not fully compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17% (only async has a description). The tool description lists four out of six parameters and explains their role in generating recommendations, adding context beyond the schema. But it omits utilization and async, leaving gaps. Overall, partial but helpful compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes AWS/Azure pricing and RIPE demand to generate Reserved Instance recommendations. It uses specific verbs and resources, and the unique combination of cloud cost optimization and RIPE trends distinguishes it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists inputs (cloud provider, instance family, region, term) and outputs (cost savings, optimal quantity, demand insights), providing clear context for use. However, it does not specify when not to use this tool or mention alternative tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

code_review_depth_optimizerA
Read-onlyIdempotent
Inspect

As a CTO, this tool analyzes your team's historical DORA metrics (deployment frequency, lead time, MTTR, change failure rate) and GitHub pull request data to recommend an optimal code review depth. Input your repository identifier and time range, and receive a structured recommendation on review rigor (light, standard, thorough) with supporting metrics and risk-adjusted rationale.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
teamSizeNoNumber of active developers in the team
repositoryYesGitHub repository identifier in format owner/repo
riskToleranceNoOrganization's risk tolerance level
timeRangeDaysYesNumber of days of historical data to analyze

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
recommendationNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds context about analyzing historical data, which aligns with read-only behavior, but does not provide additional behavioral details beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with purpose, and contains no fluff. Each sentence adds value, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description does not need to explain return values. It covers the main inputs and outputs adequately, though it could briefly mention optional parameters (teamSize, riskTolerance) that fine-tune the recommendation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description mentions 'repository identifier and time range' but adds no extra semantics beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('analyzes...to recommend') and clearly identifies the resource (code review depth). It distinguishes itself from sibling tools like 'dora_metrics_deep_dive' by focusing on a recommendation rather than pure metrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for setting code review depth based on DORA and PR data, but does not explicitly state when to use it versus alternatives (e.g., when not to use it). No exclusions or alternative tools are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

comp_benchmark_geo_deltaA
Read-onlyIdempotent
Inspect

Compares local compensation benchmarks against HQ standards for CHROs, adjusting for cost-of-living and tax differentials. Inputs include job role, local and HQ locations, and salary range. Outputs include adjusted benchmark delta, cost-of-living multiplier, and tax impact. Keywords: compensation benchmark, geographic pay equity, cost-of-living adjustment, tax differential analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
jobRoleYesStandardized job role (e.g., 'Software Engineer III')
currencyNoISO 4217 currency code (e.g., 'USD')
baseSalaryNoCurrent base salary in local currency
hqLocationYesHQ location (ISO 3166-2 code or city, country)
localLocationYesLocal work location (ISO 3166-2 code or city, country)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
taxImpactNoEstimated tax differential percentage
adjustedSalaryNoSalary adjusted for cost-of-living and taxes
benchmarkDeltaNoPercentage difference between local and HQ benchmark
confidenceScoreNo0-1 confidence in data quality
costOfLivingMultiplierNoLocal cost-of-living index relative to HQ
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly, openWorld, and idempotent hints. The description adds valuable behavioral context: the tool adjusts for cost-of-living and tax differentials and outputs specific metrics like adjusted benchmark delta, cost-of-living multiplier, and tax impact. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with the main action, and structured in three sentences. Every sentence adds value, including a list of keywords. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description adequately covers inputs and outputs. It mentions the main outputs (adjusted benchmark delta, cost-of-living multiplier, tax impact). Could be slightly improved by noting when not to use it, but overall complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and all parameters have descriptions. The description reiterates key inputs (job role, locations, salary range) but does not add significant new meaning beyond the schema definitions. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: comparing local compensation benchmarks against HQ standards for CHROs, with specific inputs and outputs. It distinguishes itself from sibling tools like executive_comp_peer_benchmark and global_salary_inflation_adjuster by focusing on geographic pay equity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for geographic compensation benchmarking but does not explicitly provide when-to-use vs when-not-to-use or alternative tools. It includes keywords that help, but lacks explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitive_deep_diveA
Read-only
Inspect

Gold-standard competitive deep dive — STRUCTURED multi-source data (no LLM narrative). Pair tool: competitor_intel for LLM-narrated board briefing + slide script. Aggregates Wikipedia, Yahoo Finance, SEC EDGAR, Wayback Machine, DuckDuckGo, HackerNews, domain scraping — all keyless. Returns agent-shaped JSON: KPIs (funding, employees, revenue, market cap), P0/P1/P2 competitive signals, pricing radar, competitor comparison matrix, Wayback timeline, positioning (sector/industry/icp_hypothesis/moat_signals), quality score. Every field is sourced or marked unavailable — no hallucinated figures. SLA: p50 ~25s, p95 ~30s · score 80+ on listed targets (US/EU/foreign) · score ~40 on private companies (no EDGAR/Yahoo data). Use sync for batch agents (≤30s tolerance). Use competitive_deep_dive_async + competitive_deep_dive_result(job_id) for conversational agents. Inputs: company name or domain (required), optional competitor list (≤5), optional depth (easy/medium/hard).

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
depthNoResearch depth: 'easy' = Wikipedia + DDG (fast, ~15s); 'medium' = + Yahoo Finance + EDGAR + Wayback (default, ~45s); 'hard' = + HackerNews + domain surfaces + competitor deep dive (~120s)
companyYesName or domain of the target company (e.g. 'Salesforce', 'notion.so', 'HubSpot CRM')
competitorsNoOptional list of competitor names or domains to include in the comparison matrix (max 5)

Output Schema

ParametersJSON Schema
NameRequiredDescription
kpisYesKey Performance Indicators sourced from public data
companyYes
qualityYes
signalsYesCompetitive intelligence signals, severity-ranked P0 (critical) to P2 (informational)
sourcesYes
comparisonYesFeature/dimension comparison between target and each competitor
depth_usedYes
positioningYesPositioning analysis derived from public data
generated_atYes
pricing_radarYesPricing tiers extracted from public sources
domain_resolvedYes
wayback_timelineYesHistorical snapshots of the company website from Wayback Machine
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description transparently discloses that the tool returns structured JSON with sourced fields, no hallucinated figures, and provides SLA (p50 ~25s, p95 ~30s) and performance expectations for different company types. This adds significant behavioral context beyond the annotations, which already indicate read-only and non-destructive behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively long but well-structured and front-loaded. Every sentence adds value, covering purpose, data sources, output format, SLA, and usage guidance. It could be slightly more condensed, but overall it is efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multiple data sources, optional depth, async option, output structure) and the presence of an output schema, the description is highly complete. It explains what the tool does, when to use it, performance characteristics, and limitations (e.g., lower scores for private companies).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the baseline is 3. The description adds additional context by explaining the async parameter's use cases (sync for batch, async for conversational) and reiterating the depth levels, which map closely to schema descriptions but are reinforced in a practical context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs a structured competitive deep dive using multi-source data with no LLM narrative. It distinguishes itself from the sibling tool 'competitor_intel' which provides an LLM-narrated board briefing. The purpose is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides usage guidance: it recommends using the synchronous version for batch agents with ≤30s tolerance and the async version for conversational agents. It also pairs with 'competitor_intel' for narrated briefings, giving clear context on when to use alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitive_deep_dive_asyncA
Read-only
Inspect

Async variant of competitive_deep_dive. Returns immediately (<200ms) with a job_id. The research runs in the background (p50≈25s, p95≈30s for depth=medium). Poll the result with competitive_deep_dive_result(job_id) after the eta_seconds hint. Use this instead of competitive_deep_dive when the agent cannot wait >15s for a response. Inputs: same as competitive_deep_dive — company (required), competitors (optional list, max 5), depth (easy/medium/hard, default medium). Async tool — register a webhook via webhooks_manage(register, url, [job.completed]) to receive callbacks instead of polling. Faster + lighter.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoResearch depth: 'easy'≈15s, 'medium'≈30s (default), 'hard'≈60s
companyYesName or domain of the target company (e.g. 'Salesforce', 'notion.so')
competitorsNoOptional list of competitor names or domains to include in the comparison matrix (max 5)

Output Schema

ParametersJSON Schema
NameRequiredDescription
job_idYesUnique job identifier — pass to competitive_deep_dive_result
statusYesAlways 'queued' on submission
eta_secondsYesEstimated seconds until result is ready
submitted_atYesISO-8601 submission timestamp
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses async behavior (returns immediately, background run time p50≈25s, p95≈30s), polling via job_id, and webhook registration. Annotations are consistent (readOnlyHint, not destructive), and description adds value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured: starts with purpose, then behavior, then input summary. Slightly verbose but all information is relevant and earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Completely covers async behavior, return value (job_id), polling mechanism, and webhook alternative. With an output schema present, return values are clear. Sibling tool for result retrieval is mentioned.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline 3. Description reiterates parameter meanings with defaults, but adds no new information beyond the schema's existing descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is an async variant of competitive_deep_dive, returning a job_id immediately. It specifies inputs and distinguishes from the sync version and the result polling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this instead of competitive_deep_dive when the agent cannot wait >15s for a response.' It also mentions webhook callback as an alternative to polling, providing clear usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitive_deep_dive_resultA
Read-onlyIdempotent
Inspect

Poll the result of a competitive_deep_dive_async job. Returns status=pending while running, status=completed with the full report once done, status=failed on error, or status=not_found if the job_id is unknown or expired (TTL 24h). Call this after the eta_seconds hint returned by competitive_deep_dive_async.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by competitive_deep_dive_async

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, idempotent, non-destructive. The description adds useful behavioral details: possible statuses, TTL of 24h, and that it returns a report upon completion. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very concise, covering purpose, statuses, TTL, and usage instruction in a few sentences. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and the presence of an output schema, the description provides all necessary context: what it does, possible statuses, TTL, and when to call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (job_id) with schema coverage 100%. The description does not add additional meaning beyond the schema's description. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it polls the result of a competitive_deep_dive_async job and enumerates possible statuses (pending, completed, failed, not_found). It distinguishes itself from sibling tools like competitive_deep_dive and competitive_deep_dive_async by focusing on result retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells agents to call this after the eta_seconds hint from competitive_deep_dive_async, providing clear usage context. However, it does not explicitly mention when not to use or alternative tools for error handling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitor_intelA
Read-onlyIdempotent
Inspect

LLM-narrated competitive-intelligence BRIEFING — for human consumption (board meeting, pitch prep). Pair tool: competitive_deep_dive for raw structured multi-source data (agent-shaped JSON). Returns: recent competitor moves with severity (critical/high/medium/low), prioritised signals, pricing-radar comparison, 3-6 quantified recommendations (impact in € or %, 7/30/90/180-day horizons), and an 8-12 slide presenter script. Use when the buyer wants a narrative briefing or a deck. Inputs: your company (name + one-paragraph pitch) + 1-10 competitors. Delivered by Manue, AI CMO of the Gapup portfolio.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNoOptional — what the buyer wants to track first (e.g. pricing moves, hiring patterns)
competitorsYes1-10 competitors to analyze
selfCompanyYesYour company info

Output Schema

ParametersJSON Schema
NameRequiredDescription
kpisNo3-5 headline KPI bubbles
sourcesNoCited sources
pricingRadarNoPricing comparison across competitors
competitorMovesYesRecent moves per competitor with severity rating
presenterScriptYes8-12 slide board presenter script
recommendationsYes3-6 actionable strategic recommendations
executiveSummaryYesBoard-ready prose summary (120-400 chars)
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false. Description adds output structure (severity, signals, recommendations, script) but does not disclose additional behavioral traits beyond annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is informative and front-loaded with purpose and output, though slightly lengthy. Every sentence adds value, no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 params, nested objects, and output schema, the description covers inputs, output format, use cases, and even persona. Return values are explained with sufficient detail given the output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all parameters. Description repeats input structure (company + competitors) and mentions async/focus, but adds minimal new meaning beyond the schema. Baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool produces an LLM-narrated competitive intelligence briefing for human consumption, distinguishing it from the sibling 'competitive_deep_dive' which provides raw structured data. Verb 'briefing' and resource 'competitive intelligence' are specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use when the buyer wants a narrative briefing or a deck' and pairs with 'competitive_deep_dive' for raw data, providing clear when-to-use and alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitor_movesC
Read-only
Inspect

Mouvements concurrents — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Answers: What have my named competitors done recently — releases, pricing changes, hires, funding? · Which competitor signals are the most urgent right now and what should I do about them? Reference case: Notion — moves de ClickUp, Asana, Coda. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
competitorsYes
selfCompanyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true (safe read) and openWorldHint=true (external dependencies). The description adds 'Returns a structured, audited deliverable' and 'Inputs are validated server-side', but omits details like potential rate limits or data freshness. With good annotations, the description adds moderate behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately concise but includes a French intro and a reference case that adds length without critical guidance. The core information is front-loaded (purpose and answers) but could be tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low schema coverage, absence of output schema, and nested object parameters, the description is incomplete. It does not clarify the expected structure of the deliverable, required subfields for selfCompany/competitors, or how to use focus. The tool is complex enough that more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (only async described). The description mentions 'named competitors' but does not explain selfCompany, competitors, or focus parameters. It refers to 'send the documented case fields' without specifying them. The description fails to compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool outputs a structured deliverable about competitor moves (releases, pricing, hires, funding) and urgent signals. The verb 'Gapup agent-payable C-suite expertise (CMO)' is somewhat opaque but the core purpose is identifiable. It does not explicitly differentiate from siblings like competitor_intel or competitive_deep_dive, but the focus on 'moves' provides some distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use vs. alternatives (e.g., competitor_intel, competitor_profiles). No when-not or exclusion criteria. The reference case (Notion vs. ClickUp, Asana, Coda) illustrates usage but does not provide decision rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitor_pricing_radarB
Read-only
Inspect

Radar pricing concurrents — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Answers: How do my competitors' pricing plans and monthly prices compare to mine? · Which competitor plan undercuts or out-features my equivalent tier? Reference case: Notion — pricing vs ClickUp, Asana, Coda. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
competitorsYes
selfCompanyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true. The description adds that it returns a 'structured, audited deliverable' and mentions server-side validation. It does not detail other behaviors like data freshness, rate limits, or what 'open world' implies. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately sized but includes some redundancy (e.g., the title appears in both annotations and description). The key purpose is front-loaded, and the reference case is helpful. However, it could be more concise and better structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested objects, 4 parameters, no output schema), the description is incomplete. It does not explain the return format, typical response time, or how the 'focus' parameter affects results. The structured deliverables are not detailed, leaving agents without enough context for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (25%). The description mentions 'selfCompany' and 'competitors' but does not elaborate on 'focus' or 'async' parameters. It adds minimal meaning beyond the schema, failing to compensate for the low coverage. The phrase 'send the documented case fields' is vague.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: comparing competitor pricing plans and monthly prices. It uses a specific verb ('Radar') and refers to a concrete reference case (Notion vs ClickUp, Asana, Coda). While 'Radar pricing concurrents' is somewhat hybrid, the overall intent is clear and distinguishes it from sibling tools like competitor_pricing_scrape.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a structured deliverable comparing competitor pricing is needed, and it mentions that inputs are validated server-side. However, it does not explicitly state when to use this tool over alternatives (e.g., competitor_profiles or competitor_pricing_scrape) or provide exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitor_pricing_scrapeA
Read-only
Inspect

Scrape and parse a competitor pricing page from a URL or domain. Fetches via proxy-aware timedFetch (tries /pricing, /plans, homepage fallback), then extracts: plan names, prices, billing cadence (monthly/annual/usage-based/one-time), key features, free tier presence, enterprise tier, estimated price range. Returns structured pricing tiers. If unfetchable or no pricing found (anti-bot, SPA, auth wall): returns a clear degraded result with warnings and signals — never fake success. ICP: founders, product managers, pricing strategists, competitive intel teams. Proxy-aware (AICI_RESEARCH_PROXY_URL). Cache TTL 6h.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesCompetitor URL or domain (e.g. 'https://notion.so/pricing', 'notion.so', 'https://www.example.com'). For best results, provide the direct pricing page URL.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.

Output Schema

ParametersJSON Schema
NameRequiredDescription
tiersYes
domainYes
statusYes
warningsYes
url_fetchedYes
has_free_tierYes
pricing_foundYes
quality_scoreYes
raw_price_signalsYes
has_enterprise_tierYes
plan_names_detectedYes
billing_model_signalsYes
estimated_price_rangeYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds rich behavioral context: proxy-aware fetching, URL fallback logic (/pricing, /plans, homepage), degraded result handling (never fake success), and cache TTL (6h). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured, front-loading the core action and then adding details on edge cases, audience, and technical notes. It is somewhat lengthy but every sentence contributes meaningful information. Could be slightly tighter but still effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (web scraping with fallback logic, caching, degraded results, output schema present), the description covers all key aspects: what it does, how it handles errors, audience, and technical constraints. It is complete for agent selection and invocation without requiring additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (both url and async parameters described). The tool description adds a usage hint for url ('For best results, provide the direct pricing page URL.'), but does not elaborate on async beyond what's in schema. Baseline 3 applies; minimal extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool scrapes and parses competitor pricing pages from a URL or domain. It specifies the extraction fields (plan names, prices, etc.) and the output format. However, it does not explicitly distinguish itself from sibling tools like competitor_pricing_radar, which may have overlapping functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description provides an ICP (founders, product managers, etc.) and hints at best use (providing direct pricing page URL). It implies use for scraping specific competitor pages but lacks explicit when-not or alternative tool guidance (e.g., when to use competitor_pricing_radar).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitor_profilesA
Read-only
Inspect

Profils concurrents — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Answers: What are the strengths, weaknesses and positioning of each of my competitors? · Give me a SWOT-style profile of a named competitor. Reference case: Notion — profils de ClickUp, Asana, Coda. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
competitorsYes
selfCompanyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true and openWorldHint=true. The description adds that the tool returns a structured, audited deliverable and that inputs are validated server-side, which is helpful but not extensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise but includes French and a reference case that may not be necessary. It front-loads the purpose but could be more streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested objects, no output schema), the description is moderately complete. It explains what the tool does and the type of output, but lacks details on the output structure or async behavior beyond the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only async has a description). The description does not elaborate on the parameters (focus, competitors, selfCompany) beyond mentioning 'documented case fields', failing to compensate for low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool produces competitor profiles answering questions about strengths, weaknesses, and positioning. It distinguishes from siblings like competitor_intel and competitor_moves by specifying a SWOT-style deliverable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly mentions when to use (e.g., 'Answers: What are the strengths... Give me a SWOT-style profile'). No explicit when-not instructions or alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

competitor_recommendationsC
Read-only
Inspect

Recommandations concurrentielles — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Answers: Given my competitors, what strategic actions should I take and in what order? · What should my 7/30/90/180-day competitive response plan look like? Reference case: Notion — actions face à ClickUp, Asana, Coda. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
competitorsYes
selfCompanyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states it returns a structured audited deliverable and that inputs are validated server-side, adding context beyond the readOnlyHint annotation. However, it does not disclose performance traits, data sources, or potential limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description contains useful information but is somewhat verbose with a mixed-language title and an example. It could be streamlined to prioritize key points without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema, the description should specify the deliverable format or structure. It merely says 'structured, audited deliverable' without details. The async parameter is not acknowledged in the description, leaving a gap in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only async described). The description fails to explain the required parameters selfCompany and competitors beyond a vague reference to 'documented case fields.' No additional meaning is provided for focus or the nested object fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides strategic competitive recommendations with a structured deliverable, specifying the output includes a response plan with priorities. However, it does not differentiate from sibling tools like competitive_deep_dive or battle_plan, which have overlapping purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., competitor_intel, battle_plan). The description implies it is for CMO-level strategic planning, but lacks explicit when-to-use or when-not-to-use criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

comp_plan_architectC
Read-only
Inspect

Architecture plan de commissionnement — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub — Comp Plan 8 rôles commerciaux · OTE €65-280k · Budget comp €2.1M · Quota coverage 3.2×. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
targetsYes
geographyNo
salesTeamYes
currentChallengesYes
preferredStructureNo
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description claims the tool returns a 'structured, audited deliverable' and that inputs are validated server-side. With readOnlyHint=true, no mutation is implied. However, it does not detail performance, response format, or how the async parameter affects behavior. No contradictions with annotations, but minimal additional context beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences plus a reference line) but includes jargon ('Gapup agent-payable C-suite expertise (CRO)') that may confuse. The reference case is illustrative but takes space. It could be more direct without losing utility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, nested objects, no output schema, and async option), the description is incomplete. It does not explain the deliverable structure, how to interpret results, or how filtering options (like geography, preferredStructure) affect output. The lack of output schema and low schema coverage exacerbate this deficiency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% (only 'async' has a description). The tool's description does not explain any parameter purpose beyond 'send the documented case fields'. For a tool with 7 parameters including nested objects (company, targets, salesTeam), this is insufficient. The description adds no parameter-specific semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The tool creates a commission plan architecture, as indicated by the title 'Architecture plan de commissionnement' and the mention of returning a structured deliverable. The reference case provides context. However, the description lacks an explicit verb like 'design' or 'build', and the mixed French/English wording may cause ambiguity. It distinguishes from siblings like 'executive_comp_peer_benchmark' by focusing on plan architecture rather than benchmarking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description provides a reference case but does not mention similar tools like 'comp_benchmark_geo_delta' or 'executive_comp_peer_benchmark'. It implies usage for designing compensation plans but offers no exclusions or context for choosing between tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_audience_profileA
Read-only
Inspect

Return the audience targeting profile of a content entity — its enrichment tags reframed as audience facets with confidence, corroboration and full provenance (verifiable, sourced). The response also carries an entity-level provenance block (average confidence, data freshness). When to use this tool: an ad-tech or marketing agent needs a machine-readable, verifiable audience descriptor for a franchise or work. Input: an entity_id and its type.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
entity_idYesEntity id from content_catalog
entity_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
entity_idYes
provenanceYesEntity-level trust & freshness summary.
entity_typeYes
audience_facetsYesMap facet → array of { label, confidence, corroboration, source_count, sources }
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description details behavioral traits: returns confidence, corroboration, full provenance (verifiable, sourced), and entity-level provenance block with average confidence and data freshness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is 4 sentences, front-loaded with the core purpose. Every sentence adds value with no waste, making it efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and annotations, the description adequately covers the tool's purpose and return value. Minor gaps like edge cases or error conditions are acceptable for a read-only tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description adds meaning by stating inputs are entity_id and its type, implying entity_type usage. It doesn't explain the async parameter but adds context beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the audience targeting profile of a content entity, with specific verb 'return' and resource 'audience targeting profile'. It distinguishes from sibling tools like content_enrichment by mentioning reframing enrichment tags into audience facets with confidence and provenance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit 'When to use this tool' section specifying the context for ad-tech or marketing agents needing verifiable audience descriptors. While it doesn't explicitly state when not to use, it provides clear context and input requirements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_catalogA
Read-only
Inspect

Browse the Gapup gold-standard content catalogue — video games, films, TV series and music. Returns franchises with their works (title, release year). When to use this tool: an agent needs structured, audited metadata for a cultural franchise, wants to resolve a title to a canonical entity, or browses a domain's catalogue before requesting enrichment. Inputs: a content domain and an optional case-insensitive name filter. Each franchise id can be passed to content_enrichment for its fine-grained tag profile.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional case-insensitive substring filter on franchise name
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum number of franchises to return (default 20)
domainYesContent domain to browse

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
domainYes
franchisesYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the agent knows it's a safe read operation. The description adds behavioral details: returns franchise works with title and release year, supports async polling via async parameter, and name filter is case-insensitive. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, well-structured, and front-loaded. It opens with the core purpose, followed by return info, usage guidelines, inputs, and relationship to another tool. Every sentence adds value with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema, the description adequately covers the return structure (franchises with works, title, release year) and links to content_enrichment. It provides complete guidance for a browsing tool, including usage context and input requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description reiterates inputs ('a content domain and an optional case-insensitive name filter') but adds no new meaning beyond the schema. The description does not add significant parameter-level detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Browse the Gapup gold-standard content catalogue — video games, films, TV series and music. Returns franchises with their works (title, release year).' This provides a specific verb+resource and distinguishes it from sibling tools like content_enrichment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'When to use this tool' section: 'an agent needs structured, audited metadata for a cultural franchise, wants to resolve a title to a canonical entity, or browses a domain's catalogue before requesting enrichment.' This gives clear context but doesn't mention when not to use or alternatives beyond content_enrichment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_compareA
Read-only
Inspect

Compare the tag profiles of two content entities (franchises or works) and measure how similar they are. Returns a Jaccard similarity score, the list of shared tags, the tags unique to each entity, and a breakdown of shared tags by facet. When to use this tool: an agent needs to compare two franchises or works (e.g. 'how similar are Dark Souls and Elden Ring?', 'what do Street Fighter and Mortal Kombat have in common?', 'on which axes do these two games differ?'), find positioning overlap, identify cross-sell opportunities, or answer 'if you liked X you might like Y' questions backed by data. Works for any domain (video-games, music, film, tv).

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
entity_aYesId of the first entity from content_catalog (e.g. 'game-dark-souls', 'music-daft-punk').
entity_bYesId of the second entity from content_catalog (e.g. 'game-elden-ring', 'music-justice').
entity_typeNoWhether both ids are franchises or works (applies to both). Defaults to 'franchise'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
entity_aYes
entity_bYes
similarityYesJaccard index = |shared| / |union|, rounded to 2 decimal places. 0 = no overlap, 1 = identical profiles.
a_tag_countYes
b_tag_countYes
entity_typeYes
shared_tagsYesTags present in both entities (up to 40).
unique_to_aYesTags present only in entity_a (up to 40).
unique_to_bYesTags present only in entity_b (up to 40).
shared_countYes
shared_by_facetYesCount of shared tags per facet (e.g. { genre: 3, theme: 5 }). Shows which dimensions drive the similarity.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and openWorldHint=true. Description details the return structure (Jaccard, shared tags, etc.) and explains async behavior. No contradictions; adds behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single concise paragraph with clear separation of purpose and usage guidelines. No wasted sentences, though slightly longer than minimal. Well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Comprehensive coverage: purpose, output, usage scenarios, parameter context, and domain generality. Output schema exists but description still explains return values, making it self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for each parameter. Description adds value by giving concrete ID examples (e.g., 'game-dark-souls') and explaining async parameter usage with polling reference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool compares tag profiles of two content entities and measures similarity, listing specific outputs (Jaccard score, shared/unique tags, facet breakdown). It distinguishes from siblings like content_similar by specifying comparison of two entities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'When to use' section with concrete examples (comparing franchises/works, finding overlap, cross-sell, recommendation) and states domain agnosticism. Lacks explicit when-not-to-use or alternative tool names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_discoveryA
Read-only
Inspect

Discover content franchises within a domain. Two modes: pass tag for a precise taxonomy match (every game tagged 'co-op'), or pass query for free-text SEMANTIC search powered by pgvector embeddings — finding franchises by meaning ('dark atmospheric games about isolation') even when no literal tag matches. Results are verifiable: tag mode carries tag confidence/corroboration, semantic mode carries a similarity score; both carry entity freshness. When to use: an agent wants a domain-scoped shortlist by tag or by intent. Inputs: a domain plus either a tag or a free-text query.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoTag label to match precisely (e.g. 'thriller', 'co-op'). Mutually exclusive with `query`.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum franchises to return (default 25)
queryNoFree-text intent for semantic search (e.g. 'melancholic synth-pop about heartbreak'). Mutually exclusive with `tag`.
domainYesContent domain to search within

Output Schema

ParametersJSON Schema
NameRequiredDescription
tagNo
countYes
queryNo
domainYes
methodYes
franchisesYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description adds details about result verifiability (confidence, corroboration, similarity scores, freshness) and the async mode. This gives agents a good sense of behavior without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, well-structured, and front-loaded with purpose. Every sentence adds value, with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 params, output schema exists), the description covers all essential aspects: purpose, modes, inputs, outputs, usage, and async behavior. Nothing important is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the description still adds significant value: explaining the mutual exclusivity of tag/query, how each mode works, and what results contain. This goes well beyond the schema's mechanical descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool discovers content franchises in a domain with two distinct modes (tag and query). It is specific about the resource and actions, but does not explicitly differentiate from sibling tools like content_catalog or content_ranking, which may have overlapping functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when to use each mode ('pass tag for precise match', 'pass query for semantic search') and a general usage statement. However, it lacks when-not-to-use guidance or comparisons to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_engineD
Read-only
Inspect

Moteur de contenu — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Notion — content engine 2026 (productivity B2B). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
brandYes
monthsYes
clusterYes
maxArticlesPerMonthYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side, but overall it contributes minimal behavioral context beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but poorly structured, mixing French and English without a clear purpose. It is front-loaded with jargon rather than actionable guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested objects, 5 parameters, no output schema), the description is insufficient. It fails to explain what the deliverable contains, how to format inputs, or what the agent should expect as a result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, and the description does not clarify the meaning or usage of parameters like brand, cluster, months, or maxArticlesPerMonth. It only vaguely refers to 'documented case fields' without elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says 'Moteur de contenu' and 'Returns a structured, audited deliverable,' but it lacks a specific verb and resource. It references a case example but does not clearly state what the tool does compared to sibling content tools like content_catalog or content_ranking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description mentions a reference case but does not explain when to choose content_engine over other content tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_enrichmentA
Read-only
Inspect

Return the enriched tag profile of a content entity — the Gapup moat. Each tag carries a facet (genre, theme, play-mode, perspective…), a confidence score, a corroboration score and its full provenance (which sources corroborated it, when). The response also carries an entity-level provenance block (average confidence, data freshness). When to use this tool: an agent has a franchise or work id (from content_catalog) and needs a fine-grained, machine-readable, verifiable characterisation for matching, recommendation, contextual targeting or analysis. Inputs: an entity id and its type.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
entity_idYesEntity id from content_catalog (e.g. 'music-daft-punk', 'film-the-dark-knight-collection:the-dark-knight')
entity_typeNoWhether the id is a franchise or a work (default franchise)

Output Schema

ParametersJSON Schema
NameRequiredDescription
tagsYes
entity_idYes
tag_countYes
provenanceYesEntity-level trust & freshness summary.
entity_typeNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. The description adds behavioral context by detailing the output structure (tag profile with facet, scores, provenance) and entity-level provenance. No contradictions; it enhances transparency beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise (3–4 sentences) with front-loaded purpose. It efficiently covers purpose, output structure, use case, and inputs without redundancy. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters and an output schema, the description is complete. It explains the tool's purpose, what the output contains, when to use it, and inputs. The existence of an output schema reduces the need to detail every field.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so description adds limited value beyond schema descriptions. It mentions entity_id comes from content_catalog and notes entity_type default, but these are already in schema. No additional parameter semantics are provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it returns an enriched tag profile of a content entity with specific components (facet, confidence, corroboration, provenance). It explicitly mentions the use case (having a franchise/work id from content_catalog) and differentiates from siblings by focusing on fine-grained, machine-readable characterization.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'When to use this tool' section provides specific conditions (entity id from content_catalog, need for characterization) and lists example applications. It does not explicitly state when not to use or name alternatives, but the guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_evergreen_score_analyzerA
Read-onlyIdempotent
Inspect

Evaluates content evergreen potential for CMOs by analyzing historical traffic patterns and backlink authority. Takes a content URL and optional time range, returns an evergreen score (0-100), traffic trend analysis, and backlink profile. Ideal for content strategy planning, SEO optimization, and identifying high-value evergreen assets. Uses Wayback Machine and Common Crawl public APIs.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesContent URL to analyze
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
toDateNoEnd date for historical analysis (YYYY-MM-DD)
fromDateNoStart date for historical analysis (YYYY-MM-DD)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
lastSeenNo
warningsYes
firstSeenNo
trafficTrendYes
backlinkCountNo
evergreenScoreYes
backlinkDomainsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly and idempotent. Description adds that the tool uses Wayback Machine and Common Crawl public APIs, and mentions async behavior via the async parameter. This provides useful context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with clear front-loading: first sentence states core purpose, second lists inputs/outputs, third adds use cases and implementation details. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists (not shown but known), the description adequately covers purpose, inputs, outputs, and external dependencies. No critical missing information for an agent to decide usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the schema already describes all parameters. The description only restates 'URL and optional time range' without adding new details like format constraints or relationship between parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool evaluates content evergreen potential for CMOs using historical traffic and backlink authority. It specifies inputs (URL, optional time range) and outputs (score, trends, backlinks). This distinguishes it from sibling tools like content_audience_profile or content_discovery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description suggests ideal use cases (content strategy planning, SEO optimization, identifying evergreen assets). It does not provide explicit when-not-to-use or alternative tools, but the context is clear enough for an AI agent to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_provenanceA
Read-only
Inspect

Audit the full data provenance of a content entity — all its enrichment tags with their extraction source, corroboration score, source list and last verification date, plus an entity-level freshness summary. Use this tool before citing or relying on enriched content data in a high-stakes context (ad targeting, editorial, analysis). Inputs: entity_id (required) and entity_type (franchise or work).

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
entity_idYesEntity id from content_catalog (e.g. 'video-game-elden-ring')
entity_typeNoWhether the id is a franchise or a work (default: franchise)

Output Schema

ParametersJSON Schema
NameRequiredDescription
lineageYesFull tag lineage from v_data_lineage — one entry per tag.
entity_idYes
entity_typeYes
freshness_summaryYesEntity-level freshness & trust summary from v_entity_freshness.
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the tool is safe and result variability is expected. The description adds detail on the output contents (tags, scores, freshness), but does not discuss authorization needs, rate limits, or other behavioral aspects beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the tool's purpose and use case, then lists inputs. It is reasonably efficient, though the 'Inputs:' section repeats schema information slightly, adding minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and annotations covering read-only behavior, the description sufficiently explains the tool's purpose, use case, and key output elements. It is complete for the context of a provenance audit tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (async, entity_id, entity_type) with descriptions. The description redundantly lists inputs but adds no new semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool audits the full data provenance of a content entity, listing specific elements like enrichment tags, extraction source, corroboration score, and freshness summary. This specifies the verb and resource, and distinguishes from sibling content tools by focusing on provenance rather than discovery or enrichment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises using the tool before citing or relying on enriched content in high-stakes contexts (ad targeting, editorial, analysis). This provides clear when-to-use guidance, though it does not name specific alternatives or when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_rankingA
Read-only
Inspect

Return the TOP-ranked content entities in a category, by a chosen criterion — the direct answer to superlative / decision queries: 'best video games', 'top RPGs', 'cheapest games', 'best value RPGs', 'best FPS playable right now', 'most popular music artists'. Criteria: critic_score, popularity, price, value (critic score per unit price). direction flips it (asc = cheapest/lowest first). available_only restricts to entities currently buyable. Sliceable by genre and release-year window; every result carries its score, price and source. When to use: an agent must produce a ranked shortlist to support a recommendation, a purchase or a 'what is the best X' decision.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
genreNoOptional genre filter, e.g. 'RPG', 'FPS', 'thriller'
limitNoNumber of ranked results (default 20)
domainYesContent domain to rank within
year_toNoOptional latest release year
criterionNocritic_score (0-100, default) · popularity · price · value (critic score per unit price)
directionNodesc = best/highest first (default); asc = cheapest/lowest/least first. Defaults to asc for price.
year_fromNoOptional earliest release year
available_onlyNoIf true, restrict to entities currently available to buy/play (default false)

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
genreNo
domainYes
rankingYes
year_toNo
criterionYes
directionNo
year_fromNo
available_onlyNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and openWorldHint. The description adds behavioral details beyond these: it explains direction defaults, slicing by genre/year, and that results include score/price/source. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads key information. It efficiently conveys purpose, criteria, and usage without unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema and annotations, the description covers all essential aspects: purpose, parameters, usage, and behavioral nuances. It is complete for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds meaning by explaining the 'value' criterion and default direction for price, which is not fully captured in the enum descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it returns top-ranked content entities by a criterion, with specific examples like 'best video games'. It clearly distinguishes itself from sibling tools through its focus on ranking and superlative queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear 'When to use' section: when an agent needs a ranked shortlist for recommendations or decision queries. It lacks explicit exclusions or alternatives, but the context is well-defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_similarA
Read-only
Inspect

Find content entities similar to a given one. For embedded franchises this uses SEMANTIC vector similarity (pgvector) over the enrichment profile — surfacing entities that feel alike even when their tags differ literally. Falls back to shared enrichment-tag overlap for works or non-embedded entities. Each result carries a similarity score and its entity-level freshness/confidence (verifiable, sourced). When to use this tool: an agent wants recommendations or lookalikes for a franchise or work. Input: an entity_id and its type.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNo
entity_idYesEntity id from content_catalog
entity_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
methodYesHow similarity was computed.
similarYes
entity_idYes
source_provenanceYesProvenance of the source entity used to compute similarity.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. The description adds significant detail: uses pgvector semantic similarity for embedded franchises, falls back to tag overlap for works, and describes result contents (similarity score, freshness/confidence). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the main action, uses four concise sentences with no redundancy. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema, the description adequately covers purpose, mechanism, fallback, and result contents. It is complete for a similarity tool with read-only and open-world hints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, and the description adds value by specifying that entity_id requires a type (entity_type is implicit). It does not describe async or limit beyond schema, but the schema covers async well. Output description does not directly add parameter semantics but clarifies expected input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Find content entities similar to a given one' with specific verb and resource. It distinguishes the mechanism (semantic vs tag overlap) and input requirements (entity_id and type). It is distinct from sibling tools like content_catalog or content_discovery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'When to use this tool: an agent wants recommendations or lookalikes for a franchise or work.' This is clear usage guidance. It does not explicitly mention when not to use alternatives, but the context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

content_taxonomyA
Read-only
Inspect

Return the enrichment taxonomy of a content domain — every tag grouped by facet (genre, theme, mood, play-mode…). When to use this tool: an agent needs the controlled vocabulary to filter, classify or query content. Input: a domain.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
domainYesContent domain

Output Schema

ParametersJSON Schema
NameRequiredDescription
domainYes
taxonomyYesMap facet → array of tag labels
tag_countYes
facet_countYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that it returns a taxonomy but does not elaborate on additional behavioral traits beyond what the annotations provide. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, concise and front-loaded with the purpose. Every sentence adds value: the first explains the result, the second explains when to use it. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 params, enum domain, output schema exists), the description covers purpose, usage, and input. It does not discuss error conditions or performance, but these are not critical for a straightforward read-only lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description mentions 'Input: a domain' which is redundant with the schema. It does not add meaning for the 'async' parameter, but the schema itself is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Return the enrichment taxonomy of a content domain — every tag grouped by facet (genre, theme, mood, play-mode…)', providing a specific verb+resource and distinguishing this tool from sibling tools that deal with content but not taxonomy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit 'When to use this tool: an agent needs the controlled vocabulary to filter, classify or query content.' This clearly defines the context, though it lacks explicit when-not-to-use or alternatives, which are not critical here.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contract_risk_scannerC
Read-only
Inspect

Scanner de risques contractuels — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: Salesforce MSA — revue d'un client SaaS B2B EMEA. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
contractTextYes
contractContextYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and openWorldHint. The description adds minimal behavioral context, only mentioning server-side validation and a reference case. It does not address the async parameter or expected output format beyond a vague 'deliverable'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but at the expense of clarity. It includes a cryptic phrase and does not front-load essential information. The structure is jumbled, mixing title, jargon, and example without logical flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given four parameters, no output schema, and low schema coverage, the description is critically incomplete. It fails to explain output structure, risk categories, or how to use the focus parameter, leaving the agent with insufficient information for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 25% schema description coverage, the description should explain parameter semantics but does not. It only hints at 'documented case fields' without defining the required contractContext fields or the purpose of async, focus, or contractText.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it scans contractual risks and returns a structured deliverable, providing a reference case. However, the phrase 'Gapup agent-payable C-suite expertise (RISK)' is confusing and fails to clearly specify the verb-resource relationship. It does not effectively distinguish from sibling tools like 'legal_clause_extractor'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description lacks context for when risks should be scanned or prerequisites, leaving the agent without selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

corporate_registry_lookupA
Read-onlyIdempotent
Inspect

Resolve legal information about a company from its national corporate registry. Returns a normalised, sourced company profile: legal status, registration number, directors, shareholders, recent filings, registered address, share capital, and a quality score (0–100). Coverage: France (INPI, keyless — full SIREN/SIRET with directors), 3M+ entities worldwide via GLEIF LEI (keyless, large companies), UK (Companies House, optional key), Netherlands (KvK, optional key), and OpenCorporates (token required since 2026). Sources are tried in cascade; quality_score increases with each source that succeeds. When to use: due-diligence, KYC screening, supplier verification, M&A research, or any workflow needing verified company identity and legal status. Optional env vars: COMPANIES_HOUSE_API_KEY (UK), KVK_API_KEY (NL), OPENCORPORATES_API_TOKEN (OpenCorporates token).

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
countryNoISO 3166-1 alpha-2 country code (e.g. 'FR', 'GB', 'NL', 'DE', 'SG', 'AU', 'US'). If omitted, inferred from legal suffix in company name, then falls back to global search.
identifierNoOptional registry identifier for a fast direct lookup: SIREN (FR, 9 digits), Companies House number (GB, 8 chars), KvK number (NL, 8 digits), etc.
company_nameYesCompany name or trading name to look up (e.g. 'Sanofi', 'Tesco PLC', 'Notion Labs Inc')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
registryYes
directorsYes
freshnessYesISO timestamp
identifierYes
legal_formNo
legal_nameNo
company_nameYes
jurisdictionYes
shareholdersYes
quality_scoreYes0-100 confidence score
share_capitalNo
filings_recentYes
incorporation_dateNo
registered_addressNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnly, idempotent, non-destructive. Description adds cascade logic, quality score increase, async behavior, and source-specific requirements. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured: purpose, coverage, use cases, env vars. Slightly long but no wasted sentences. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given complexity (4 params, many sources, output schema exists), description covers purpose, when-to-use, behavioral details, and parameter nuances. Output schema handles return values, so no gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 4 parameters are described in the schema (100% coverage). The description adds context: async usage, country inference, and identifier for direct lookup, enhancing understanding beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it resolves legal information from national corporate registries, lists the returned fields, and covers multiple sources. It distinguishes itself from siblings by focusing on registry lookups for due diligence and KYC.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly lists use cases (due-diligence, KYC, etc.) and mentions optional API keys. Lacks explicit when-not-to-use or alternatives, but the targeted applications are well defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

court_filings_multiA
Read-only
Inspect

Aggregate court filings, judgments and litigation records for a company or individual across five major legal jurisdictions: US (CourtListener / PACER), UK (National Archives — EWHC/EWCA/UKSC/UKUT), EU (ECHR HUDOC — European Court of Human Rights), France (Légifrance / Cour de cassation) and Germany (BGH / BVerfG). Returns structured case records with type classification (civil/criminal/antitrust/bankruptcy/administrative/unknown), status (filed/pending/decided/appealed/unknown), parties extracted from case titles, opinion URLs and verbatim snippets. Cross-case pattern recognition produces severity-ranked signals (P0–P2) for criminal, antitrust, bankruptcy, regulatory, data-breach and IP categories. Use when: due diligence on a counterparty, vendor risk assessment, competitive intelligence (litigation history), regulatory exposure mapping. All sources are public and keyless. Optional env var COURTLISTENER_API_KEY raises US rate limits beyond the default 5 req/s anonymous tier. SLA: ≤25s p95 (all jurisdictions fetched in parallel, 8s budget per source). Quality score: 20 pts per jurisdiction with ≥1 case retrieved, +10 if signals detected, +5–10 if ≥2–3 distinct sources contributed.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
date_toNoISO date YYYY-MM-DD — latest filing or decision date to include
date_fromNoISO date YYYY-MM-DD — earliest filing or decision date to include
party_nameYesName of the company or individual to search (e.g. "Apple Inc", "TotalEnergies", "Volkswagen AG")
jurisdictionNoJurisdictions to search. Defaults to all ["US","UK","EU","FR","DE"].

Output Schema

ParametersJSON Schema
NameRequiredDescription
casesYes
statusYes
signalsYes
sourcesYes
party_nameYes
quality_scoreYes
by_jurisdictionYes
jurisdictions_searchedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond annotations by detailing SLA (≤25s p95), quality scoring formula, optional rate limit key, parallel fetching strategy, and signal severity levels (P0–P2). Annotations already indicate read-only and non-destructive behavior, but the description adds operational transparency without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and informative, with clear front-loading of purpose and jurisdictions. It includes details on output, use cases, and SLA. While comprehensive, it is slightly verbose; each sentence adds value but could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (five jurisdictions, multiple data sources, cross-case pattern recognition, quality scoring) and the presence of an output schema, the description is sufficiently complete. It covers what, where, and how well the tool performs, leaving little ambiguity for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100%, so baseline is 3. The description adds context for jurisdictions (e.g., 'US (CourtListener / PACER)') but does not significantly elaborate on parameter usage beyond what the schema provides. Some parameter descriptions are in the schema already, so the added value is marginal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool aggregates court filings, judgments, and litigation records across five specific jurisdictions. It lists the data sources, what is returned, and cross-case pattern recognition. This differentiates it from sibling tools by specifying the multi-jurisdiction coverage and detailed outputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides use cases: due diligence, vendor risk assessment, competitive intelligence, regulatory exposure mapping. It notes all sources are public and keyless. However, it does not explicitly state when not to use this tool or mention specific alternative tools among the many siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crm_connectorAInspect

Push, update, search and log activities in HubSpot, Salesforce or Pipedrive. 4 modes: push_lead (create contact/lead), update_opportunity (update deal stage/amount), search_contact (lookup by email), log_activity (call/email/meeting/note). Returns resource_id, direct CRM URL, signals and quality_score. If credentials are absent, returns a mock result with a warning signal. Auth: HubSpot via Bearer access_token; Salesforce via access_token + base_url; Pipedrive via api_key.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataYesPayload depending on mode. push_lead: {email,first_name,last_name,company,phone,job_title}. update_opportunity: {deal_id/opportunity_id,stage,amount,close_date}. search_contact: {email}. log_activity: {type,body,contact_id/person_id,subject}.
modeYesAction to perform in the CRM
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
providerYesCRM provider to target
credentialsNoAuth credentials. HubSpot: access_token. Salesforce: access_token + base_url. Pipedrive: api_key.

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlNo
modeYes
statusYes
signalsYes
sourcesYes
successYes
providerYes
data_syncedNo
resource_idNo
quality_scoreYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavioral traits: it returns a mock result with a warning signal when credentials are absent, details authentication requirements per provider, and outlines return values (resource_id, CRM URL, signals, quality_score). This adds significant context beyond the annotations, which only indicate readOnlyHint=false and openWorldHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the main action. Each sentence adds value, covering modes, return values, fallback behavior, and auth. It could be more structured with bullet points, but it remains efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multiple providers, four modes, async option, auth variations), the description covers all critical aspects: mode actions, provider options, auth details, fallback behavior, return values, and async usage. It is comprehensive and leaves no obvious gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with all parameters described. The description adds a summary of payload structures per mode, which is helpful but does not provide new semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: push, update, search, and log activities in HubSpot, Salesforce, or Pipedrive. It enumerates four modes with specific actions, leaving no ambiguity about what the tool does. Given the extensive sibling list with no direct competitors, the purpose is well-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool versus alternatives. It does not mention exclusions or alternative tools for similar tasks. Usage is implied through mode descriptions but not stated, limiting the agent's ability to decide appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cross_sell_recoC
Read-only
Inspect

Recommandations cross-sell — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Alan × Gapup Hub — 3 produits recommandés · Fit 'perfect' × 2 · ARR potentiel +€18k. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
accountYes
companyYes
portfolioYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds context by stating the output is a 'structured, audited deliverable' and provides a reference case (Alan × Gapup Hub) that illustrates typical output content (products, fit, ARR potential). This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, but the second sentence (reference case) is moderately informative yet somewhat noise. The first sentence is jargon-heavy ('Gapup agent-payable C-suite expertise (CRO)'). A more concise explanation of purpose and parameters would improve.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested objects, no output schema, low schema coverage), the description falls short. It doesn't explain the return structure, required nested fields (e.g., account.currentProducts), or how the portfolio array is used. The reference case gives a glimpse but lacks systematic coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 25% schema description coverage, the description should compensate by explaining the required parameters (account, company, portfolio). Instead, it vaguely says 'send the documented case fields' without detailing what fields or how they map to the input schema. No parameter meaning is added beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it provides cross-sell recommendations via a 'structured, audited deliverable', clearly indicating the verb (recommend) and resource (cross-sell opportunities). However, it does not differentiate from the sibling tool 'upsell_hunter', which likely serves a similar purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The mention of 'Gapup agent-payable C-suite expertise (CRO)' hints at a sales audience but provides no exclusions or context that helps an agent decide between this and similar tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crypto_wallet_intelA
Read-only
Inspect

Multi-chain on-chain analytics for crypto trading agents, on-chain analysts, AML/compliance teams and DeFi BD. Covers Ethereum, Base, Polygon, BSC, Arbitrum, Optimism — EVM-compatible addresses only.

5 modes: • wallet_profile — full wallet summary: type (EOA/contract/CEX/protocol), inferred persona (whale/MEV-bot/DeFi-user/hodler…), age, tx count, native balance, ERC-20 count, NFT collections, OFAC sanctions flag • token_flows — ERC-20 inflows/outflows per token on the selected period, priced in USD via CoinGecko • pnl_estimate — FIFO realized + unrealized P&L on the period with confidence rating (high/medium/low) • counterparties — top 20 counterparties ranked by USD volume with CEX/DEX/protocol labels • defi_positions — active DeFi positions detected via Etherscan interaction history (Aave/Compound/Uniswap/Curve/Lido/Balancer/SushiSwap)

Signal detection (P0/P1/P2): P0 if OFAC SDN match OR direct Tornado Cash / sanctioned-protocol interaction P1 if >$1M volume on wallet <30 days old OR MEV-bot pattern OR >80% volume on single counterparty P2 informational (CEX wallet, new wallet, no anomaly)

Sources: Etherscan family (keyless free-tier, optional API key per chain), DefiLlama (keyless), public EVM RPC (keyless), CoinGecko free tier (keyless). Cache TTL: 5 min (wallet activity evolves fast). Budget: 8s per source.

Env vars (all optional, raise Etherscan rate-limit from 1 req/5s to 5 req/s): ETHERSCAN_API_KEY · BASESCAN_API_KEY · POLYGONSCAN_API_KEY BSCSCAN_API_KEY · ARBISCAN_API_KEY · OPTIMISM_API_KEY

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesAnalysis mode. wallet_profile=full wallet summary + persona + sanctions flag. token_flows=ERC-20 inflows/outflows per token priced in USD. pnl_estimate=FIFO realized+unrealized P&L with confidence. counterparties=top 20 counterparties by volume. defi_positions=active positions on Aave/Compound/Uniswap/Curve/Lido/etc.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
chainNoChain to analyze. Default "ethereum". Use "all" to scan all 6 chains (slower, ~30s).
addressYesEVM-compatible wallet address (0x... 40 hex chars). Works on all supported chains.
period_daysNoLookback window in days for token_flows, pnl_estimate, counterparties, defi_positions. Default 30.
min_value_usdNoMinimum USD value filter for token_flows and counterparties. Default $100.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
addressYes
signalsYes
sourcesYes
token_flowsNo
pnl_estimateNo
quality_scoreYes
counterpartiesNo
defi_positionsNo
wallet_profileNo
chains_analyzedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description details read-only nature (consistent with readOnlyHint=true), multi-source dependency, caching, performance budgets, and env vars for rate limits. It adds significant behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with an overview, bullet list of modes, signal detection, sources, caching, and env vars. Every sentence adds value without redundancy, achieving conciseness despite length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, 5 modes, signal levels, env vars), the description covers all major aspects: input details, behavioral traits, source dependencies, caching, and performance budget. The existence of an output schema (not shown) compensates for missing return value details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers all parameters with descriptions. The description adds value by explaining mode-specific details (signal detection, P0/P1/P2 levels) that aren't in the schema, though other parameters like period_days are well-documented in schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: multi-chain on-chain analytics for crypto trading agents, analysts, AML, and DeFi teams. It lists 5 distinct modes and supported chains, distinguishing it from siblings like usdc_x402_payments_intel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides extensive context on what the tool does but does not explicitly state when to use it versus alternatives. No 'use this when' or exclusion criteria are given, relying on the specificity of the tool to imply usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

customer_marketingC
Read-only
Inspect

Marketing clients & ambassadeurs — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Gapup Hub — 12 clients analysés · 4 ambassadeurs identifiés · Programme + 6 case studies + référral. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
goalsYes
companyYes
productYes
customersYes
targetUseCasesNo
contentBudgetEurNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, consistent with 'Returns a structured, audited deliverable'. The description adds that inputs are validated server-side, but does not mention async behavior or other behavioral traits. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but includes fragmented phrases and mixed languages (French/English). It could be more structured and front-loaded with a clear action statement. However, it is not overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of 7 parameters, nested objects, and no output schema, the description is insufficient. It lacks details on expected output format, behavior, and parameter relationships. The reference case helps but does not cover the full scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 14%, very low. The description does not explain any parameters, only mentioning 'the documented case fields'. It adds minimal semantic value beyond the schema, leaving the meaning of parameters like 'customers', 'goals', and 'contentBudgetEur' unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description mentions 'Marketing clients & ambassadeurs' and 'Returns a structured, audited deliverable', giving a general sense of the tool's purpose, but lacks a clear verb (e.g., 'analyze', 'generate', 'identify'). The reference case provides examples, but the primary function remains ambiguous, especially compared to sibling marketing tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not specify context, prerequisites, or when not to use it. It only states that inputs are validated server-side, which is not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

customer_voice_synthC
Read-only
Inspect

Synthèse voix client — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Alan (assurance santé) — 3 personas · Top 5 douleurs · Repositionnement messagerie recommandé. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
dataSourcesYes
targetSegmentsYes
repositioningFocusNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side and returns a structured deliverable, which is consistent. However, it does not disclose any additional behavioral traits like auth needs or rate limits, leaving the agent with minimal extra context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (3 sentences) but lacks clear structure. It mixes French and English, and the reference case detail is somewhat tangential. While concise, it is not optimally organized for an AI agent to quickly extract key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters, 3 required, no output schema, and only 20% schema coverage, the description is incomplete. It does not explain the return value, how to construct inputs beyond the reference case, or what the 'audited deliverable' contains. The agent would need to guess or rely on external knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, yet the description provides no explanation of parameters like 'company', 'dataSources', or 'targetSegments'. It only vaguely says 'send the documented case fields' without linking to schema properties. This forces the agent to rely entirely on the sparse schema descriptions, which is inadequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states it performs 'Synthèse voix client' and returns a 'structured, audited deliverable' for C-suite. This is a specific verb+resource, though the phrase 'Gapup agent-payable C-suite expertise (CMO)' adds jargon. It distinguishes from siblings by focusing on customer voice synthesis but does not explicitly differentiate from similar analysis tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention prerequisites, context, or exclusions. It only provides a reference case without clarifying when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cve_security_lookupA
Read-only
Inspect

Look up CVE vulnerability data for enterprise security teams, DevSecOps and SOC analysts. Supports two modes: exact CVE ID lookup (e.g. 'CVE-2024-3094') or keyword search by product/vendor (e.g. 'openssl', 'Apache Tomcat'). Cross-references four authoritative keyless sources: NVD NIST (official CVE database, CVSS v3 scores, affected CPEs), CISA KEV (Known Exploited Vulnerabilities catalog — exploit_in_wild flag), EPSS FIRST (exploit probability 0-1), GitHub Security Advisories (ecosystem-specific: npm/pypi/maven). Returns structured vulnerability records with CVSS v3 scores, affected product version ranges, CWE weakness classification, references and exploitation status. Signals engine produces P0/P1/P2 alerts: P0=CVSS>=9 + active exploitation, P1=CVSS>=7 or EPSS>=70%, P2=CWE pattern clusters. Relevant for EU NIS2 and DORA supply chain risk obligations. Optional env: NVD_API_KEY (raises NVD rate-limit 5→50 req/30s), GITHUB_TOKEN (raises GHSA GraphQL rate-limit). Cache TTL 6h. SLA <=25s p95.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoOverride auto-detection: "lookup" for exact CVE ID, "search" for product/vendor keyword.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesCVE ID (e.g. "CVE-2024-3094") or product/vendor keyword (e.g. "openssl", "Apache Tomcat"). Mode is auto-detected from the CVE-YYYY-XXXXX pattern.
max_resultsNoMaximum number of vulnerabilities to return (default 20, max 50).
severity_minNoMinimum CVSS v3 severity to include in results (default: no filter).
published_afterNoISO date YYYY-MM-DD — only include CVEs published after this date. Defaults to 365 days ago for search mode.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
queryYes
statusYes
signalsYes
sourcesYes
quality_scoreYes
vulnerabilitiesYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, non-destructive, open-world behavior. The description adds significant detail beyond annotations: rate-limit behavior with optional API keys, cache TTL (6h), SLA (≤25s p95), and the alerting logic (P0/P1/P2). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but well-structured, starting with the main function and progressively adding context. It could be slightly more concise (e.g., mentioning NIS2/DORA obligations might be extraneous), but every sentence contributes to understanding. Front-loaded with key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 6 parameters, an output schema, and no nested objects, the description covers all essential aspects: modes, data sources, alerting, rate limits, caching, and SLA. It is fully self-contained and leaves no ambiguity about tool behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 6 parameters have schema descriptions (100% coverage). The description adds value by explaining mode auto-detection, clarifying defaults (e.g., max_results default 20, published_after default 365 days for search), and providing concrete examples for 'query' parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: lookup CVE vulnerability data via exact CVE ID or keyword search. It specifies the target audience (enterprise security teams, DevSecOps, SOC analysts) and contrasts the two modes with examples, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use the tool (CVE lookup or keyword search) and includes examples for each mode. However, it does not explicitly differentiate from sibling tools like 'vuln_exploitability_forecast' or 'vuln_patch_priority_engine', which could be mentioned as alternatives for related but distinct tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cyber_risk_auditorC
Read-only
Inspect

Auditeur de risque cyber — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: Qonto — Audit cyber risque B2B FinTech · Score 58/100 → roadmap 90j · 8 findings critiques/high · économie prime -28%. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
techStackYes
currentPostureYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side, which is consistent. No contradictions, but additional behavior traits (e.g., what the deliverable contains) are not disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short with three sentences. It front-loads the tool's purpose. However, the reference case example adds length without providing critical usage info; a slightly more concise version would be better.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested input objects, no output schema), the description is incomplete. It does not explain output format, how to interpret results, or how to use the focus parameter. The agent lacks sufficient context to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (20%), with only the async parameter having a description. The description does not compensate by explaining the company, techStack, currentPosture, or focus parameters. It merely says 'send the documented case fields' without elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs a cyber risk audit and returns a structured deliverable. The reference case provides a concrete example. However, it does not explicitly differentiate from sibling tools, but the name and purpose are specific enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention context, prerequisites, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deal_coachB
Read-only
Inspect

Coach de deal MEDDIC — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Datadog Enterprise deal Société Générale €1.2M ARR — coaching MEDDIC + escalation plays + 14 next actions. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
dealYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
knownContextYes
buyingCommitteeYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description aligns with annotations: readOnlyHint true (returns a deliverable, no mutations), openWorldHint true (varying results). It adds behavioral context: inputs validated server-side, async option for non-blocking call. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is two sentences plus a reference case, concise and to the point. It front-loads the purpose. Could be slightly more structured (e.g., list of return fields) but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema provided; description only says 'structured, audited deliverable' without specifying contents. For a tool with 5 parameters, nested objects, and no return details, this is incomplete. Missing guidance on what to expect in the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 20% (only 'async' described). The description provides no additional detail on parameters beyond 'send the documented case fields'. The schema defines required objects (deal, buyingCommittee, knownContext) but the description does not explain their semantics or usage, leaving the agent with limited guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides MEDDIC deal coaching and returns a structured deliverable. It references a specific use case (Datadog Enterprise deal). However, it does not explicitly differentiate from similar tools like meddic_scoring or deal_structurer, though the name and focus on coaching provide some distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used for coaching deals with MEDDIC by requiring inputs like deal and buying committee. It mentions server-side validation and documented fields, but lacks explicit guidance on when to use this vs. alternatives (e.g., meddic_scoring for simpler scoring) or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deal_structurerC
Read-only
Inspect

Structuration de deal — Gapup agent-payable C-suite expertise (CSO). Returns a structured, audited deliverable. Reference case: Agicap × Kyriba — Partenariat API Banking · 5 structures comparées · Term sheet 7 clauses · Score 83/100 JV. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
dealYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true. Description adds that it returns a structured, audited deliverable and validates inputs, but no additional behavioral traits beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with some wasted space on the reference case example. The first sentence is clear but the example is verbose for a tool description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema; description should explain return value format. The reference case hints at outputs but doesn't fully describe what the deliverable contains, leaving gaps for a complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (33%). Description provides a reference case but does not explain each parameter's meaning or usage details, leaving the agent with insufficient semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb 'Structuration de deal' and resource with deliverable. Reference case adds specificity but does not differentiate from siblings like 'deal_coach' or 'term_sheet_negotiation'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. Mentions server-side validation but no context for choosing this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dependency_vulnerability_scanA
Read-only
Inspect

SCA (Software Composition Analysis) — scans a project dependency manifest and returns known vulnerabilities for each dependency. Supports: package.json (npm), requirements.txt (Python), go.mod (Go), Cargo.toml (Rust), composer.json (PHP), Gemfile.lock (Ruby), CycloneDX SBOM JSON. PRIMARY source: OSV.dev (keyless, free, covers npm/PyPI/Go/crates.io/Packagist/RubyGems + GHSA advisories federated). CVSS enrichment: NVD NIST (when OSV lacks score). Exploitation flag: CISA KEV (known-exploited-vulnerabilities catalog). Returns per-vuln CVE/GHSA IDs, severity, CVSS score, fixed version, and actionable upgrade recommendations. Relevant for EU NIS2 supply chain risk obligations, DORA, SOC 2 vendor assessments. Cache TTL 6h. Parallel OSV queries (concurrency=10). SLA <=30s p95.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesManifest type: "package_json"=npm, "requirements_txt"=pip, "go_mod"=Go modules, "cargo_toml"=Rust, "composer_json"=PHP, "gem_lock"=Ruby, "sbom_cyclonedx"=CycloneDX SBOM JSON.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
severity_minNoMinimum severity to include in results (default: "medium").
manifest_contentYesRaw text content of the manifest file to scan (e.g. full contents of package.json, requirements.txt, etc.).
include_transitiveNoInclude transitive/indirect dependencies in results (default: true).

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
sourcesYes
summaryYes
ecosystemYes
quality_scoreYes
recommendationsYes
vulnerabilitiesYes
dependencies_parsedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations by detailing data sources (OSV.dev, NVD NIST, CISA KEV), caching (TTL 6h), concurrency (10), SLA (≤30s p95), and async polling behavior. Read-only and non-destructive nature is reinforced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well-organized; front-loaded with purpose and supported formats. Every sentence contributes, though slightly long. Efficient for its information density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers input, behavior, data sources, performance, and compliance relevance. Output schema is present, so return values are documented separately. Complete for a tool with rich annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description adds extra context: enum mappings, default values, and purpose of async parameter. Adds value beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it scans dependency manifests for known vulnerabilities, lists supported manifest types, and distinguishes from general CVE lookups or other vulnerability tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage cues: specific manifest formats, compliance standards, and performance characteristics. Lacks explicit when-not or alternative recommendations, but the context is sufficient for correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discovery_prepC
Read-only
Inspect

Préparation discovery — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Discovery Salesforce × Airbus — VP Digital Marc Legrand · Signaux achat confirmés · +28 pts conversion demo. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
contactYes
ourOfferYes
prospectYes
meetingGoalNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint: true and openWorldHint: true. The description adds that it returns a structured, audited deliverable and that inputs are validated server-side. This provides some behavioral context beyond annotations but does not elaborate on what 'audited' entails or what the open world hint implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short (two sentences plus a reference case). However, the reference case is a distractor, and the language mix reduces clarity. It could be more concise and structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 parameters, nested objects, no output schema), the description is incomplete. It does not specify the return format, what the deliverable contains, or how to interpret results. The annotations provide readOnlyHint but not enough context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'async' has a description). The description does not explain individual parameters or their roles beyond implying they are 'documented case fields'. It adds minimal value beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it prepares a discovery deliverable for C-suite expertise (CRO) and returns a structured, audited output. The verb 'Préparation' and resource 'discovery' are present, and it is distinct from siblings like 'content_discovery'. However, the mixed French/English and unclear term 'Gapup' reduce clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. It only instructs to send documented case fields and mentions server-side validation, but does not specify prerequisite conditions, scenarios, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diversity_inclusion_metricsC
Read-only
Inspect

Métriques diversité & inclusion — Gapup agent-payable C-suite expertise (SUSTAINABILITY). Returns a structured, audited deliverable. Reference case: Cas démo — Métriques diversité & inclusion. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
ambitionsYes
currentStateYes
regulatoryContextNo
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds no behavioral context beyond 'returns a deliverable', and omits details like response format or potential restrictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with some redundancy (reference case and validation). The first sentence is cryptic and wastes words. Could be tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a complex nested input (6 params, 3 required) and no output schema. The description fails to explain the deliverable's structure, how to use the async parameter, or what the reference case implies. Incomplete for safe invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 17% schema description coverage, the description provides no parameter-level explanations. Terms like 'ambitions' and 'currentState' are left undefined, forcing reliance on the schema, which also lacks descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns a structured, audited deliverable for diversity & inclusion metrics, but the purpose is muddled by jargon like 'Gapup agent-payable C-suite expertise'. It does not clearly distinguish from sibling tools or specify what metrics are computed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like bias_amplification_tracker or hr_benefits_esg_aligner. The only instruction is to send documented case fields, which is generic.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

domain_tech_fingerprintA
Read-only
Inspect

Empreinte tech d'un domaine — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Answers: What is the tech stack of — frontend, CMS, analytics, CRM, CDN, hosting? · What buying signals does 's technology footprint reveal for sales prospecting? · Analyze for supply-chain technology risk and third-party vendor exposure. · What is the best outreach angle for a sales rep targeting based on their detected stack? · Run a CISO-style technology fingerprint on — identify legacy tech, missing security headers, and vendor risk. · Has recently changed their marketing or analytics stack — any vendor adoption signals? Reference case: velora-payments.io · Next.js + Cloudflare + Stripe + GA4 + HubSpot · . Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
depthYesstandard
focusYestech-buying
target_domainYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the description adds value by noting the deliverable is 'audited' and that inputs are validated server-side. It does not contradict annotations and provides context about the output's nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than needed, with a bulleted list of questions and a reference case. While front-loaded with purpose, it contains redundant details that could be trimmed. A more concise description would improve efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description adequately covers the tool's purpose and typical results through example questions. It mentions the 'async' parameter implicitly via the example reference, but does not explain its use. Overall, it is reasonably complete for a 4-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (25%), but the description partially compensates by linking the 'focus' enum values to the listed questions. However, 'depth' and 'target_domain' are not explained, and the mention of 'documented case fields' is vague. Some additional meaning is added, but not comprehensive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a structured tech stack deliverable for a domain. It lists specific questions it answers (frontend, CMS, analytics, etc.), distinguishing it from sibling tools like competitive_deep_dive or competitor_intel, which focus on broader competitive analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists use cases (buying signals, supply-chain risk, outreach angles) but does not explicitly state when not to use this tool or compare to alternatives. Usage is implied but lacks clear boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dora_metrics_deep_diveA
Read-onlyIdempotent
Inspect

Analyzes DORA metrics (Deployment Frequency, Mean Time to Recovery, Change Failure Rate) with deep correlation to code review patterns. Designed for CTOs to identify bottlenecks in software delivery pipelines. Inputs include GitHub repository identifiers and optional time ranges. Outputs structured metrics with trend analysis and code review depth insights.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoYesGitHub repository in format 'owner/repo'
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
sinceNoStart date for analysis (ISO 8601)
untilNoEnd date for analysis (ISO 8601)
branchNoBranch name to analyze (default: main)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
metricsNo
sourcesNo
warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, so safety is assured. The description adds value by stating outputs include structured metrics with trend analysis and code review depth insights, providing behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with clear structure: function, audience/purpose, inputs/outputs. Front-loaded with key information. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters and an output schema exists, the description is adequate. It covers core purpose, audience, input types, and output nature. Minor missing details on how code review correlation works are compensated by schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (all parameters described). The description only summarizes inputs as 'GitHub repository identifiers and optional time ranges', which does not add significant meaning beyond schema definitions. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes DORA metrics (Deployment Frequency, Mean Time to Recovery, Change Failure Rate) with correlation to code review patterns. It specifies the verb 'analyzes' and the resource, but does not explicitly differentiate from sibling tools like 'mttr_breakdown_analyzer' or 'code_review_depth_optimizer'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for CTOs identifying bottlenecks in software delivery pipelines, but lacks explicit guidance on when to use this tool vs alternatives, or conditions where it should not be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dora_operational_resilience_stress_tesA
Read-onlyIdempotent
Inspect

Assess DORA operational resilience by simulating ICT failure scenarios for financial entities. Designed for legal/compliance teams to evaluate ICT risk management under DORA Article 25. Inputs include failure scenario parameters (e.g., ICT service type, duration, impact radius) and entity profile. Outputs structured resilience scores, regulatory gaps, and mitigation recommendations with EUR-Lex/FTC enforcement references.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
entityTypeYes
impactRadiusYes
ictServiceTypeYes
existingMitigationsNo
failureDurationHoursYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
warningsNo
regulatoryGapsYes
resilienceScoreYes
simulationTimestampNo
recommendedMitigationsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, so the tool is understood as a safe, read-only simulation. The description adds context about outputs (resilience scores, regulatory gaps, recommendations) and inputs, but does not disclose side effects or performance characteristics beyond what annotations cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences) and front-loaded with the core purpose. Every sentence adds value, covering purpose, target audience, regulation, inputs, and outputs without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, 4 required, enums, output schema) and the presence of an output schema (which reduces the need to explain return values), the description covers most aspects: purpose, regulation, target users, and output types. It lacks elaboration on parameter details and async behavior, but these are partially covered in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17% (only the async parameter has a description). The description lists some parameter examples ('ICT service type, duration, impact radius') but does not explain the enum values or other parameters like 'existingMitigations'. This leaves significant ambiguity about parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Assess DORA operational resilience by simulating ICT failure scenarios') and identifies the target users (legal/compliance teams) and regulatory context (DORA Article 25). It distinguishes itself from sibling tools like 'dora_metrics_deep_dive' by focusing on simulation and regulatory gap analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool (for evaluating ICT risk management under DORA Article 25) and who it's designed for (legal/compliance teams). However, it does not mention when not to use it or offer alternatives, which would improve clarity further.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dual_use_export_risk_mapperA
Read-onlyIdempotent
Inspect

As a COO, quickly assess export compliance risks for components in your supply chain. This tool analyzes bills of materials (BOMs) against EU dual-use export control lists and ICAO/IMO restricted items data. Input a list of part numbers, descriptions, or HS codes to receive a risk assessment with actionable insights. Output includes risk levels, applicable regulations, and source references for audit trails.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
bomItemsYes
includeSourcesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
resultsNo
sourcesNo
warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds behavioral context: it 'quickly assesses' risks, uses EU dual-use lists and ICAO/IMO data, and produces actionable insights with audit trails. This complements the annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three sentences, front-loading the primary audience and purpose. Every sentence adds value—no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of export compliance, the description covers core functionality, inputs, and outputs. It references specific regulations (EU dual-use, ICAO/IMO) and provides output details. With an output schema presumably present, this is sufficient, though a mention of the async parameter's behavior would enhance completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (33%), but the description partially compensates by detailing that 'bomItems' is a list of items with part numbers, descriptions, or HS codes. However, 'includeSources' is not mentioned, and no further parameter guidance is given beyond the schema's limited descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool assesses export compliance risks for supply chain components, with specific inputs (BOMs, part numbers, descriptions, HS codes) and outputs (risk levels, regulations, references). However, it does not explicitly differentiate from sibling tools like 'dual_use_tech_diversion_monitor' which may share similar functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description sets a clear user persona (COO) and context (supply chain compliance), but lacks guidance on when not to use this tool or alternatives. Given siblings exist, explicit exclusions or comparison would improve this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dual_use_tech_diversion_monitorA
Read-onlyIdempotent
Inspect

Asynchronous T5-level tool for COO persona to detect unauthorized diversion of dual-use technologies. Cross-references shipment manifests, EU sanctions lists, and ICAO/IMO transport data to identify suspicious transfers. Inputs: shipment IDs, company identifiers, or geographic routes. Outputs structured diversion risk assessment with source provenance. Requires async:true to avoid 402 timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
routeNo
companyIdNoCompany registration number or tax identifier
shipmentIdNoUnique shipment identifier (e.g., bill of lading number)
techCategoryNoDual-use technology category

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
matchesNo
sourcesNo
warningsNo
diversionRiskNoCalculated diversion risk score (0-100)
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, openWorldHint, idempotentHint. Description adds valuable behavioral context: asynchronous execution, cross-referencing specific databases, output with provenance, and timeout prevention hint. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose, efficient and to the point. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, data sources, inputs, output format, and performance constraints. With output schema existing, no need to detail return structure. Missing error handling or edge cases but adequate for a compliance tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so baseline 3. Description lists input types (shipment IDs, company identifiers, geographic routes) that map to parameters, adding some context but not detailed per-parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool detects unauthorized diversion of dual-use technologies and lists data sources. However, it does not explicitly distinguish from similar sibling like 'dual_use_export_risk_mapper', leaving some overlap ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. Only mentions async requirement for performance, but no context on preferred use cases or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

earnings_reviewerC
Read-only
Inspect

Earnings Reviewer — Gapup agent-payable C-suite expertise (FUNDRAISING). Returns a structured, audited deliverable. Reference case: Salesforce Q3 FY2026 — call transcript + 10-Q + guidance → analyst note. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
quarterYes
analystFocusNo
secFilingContextNo
transcriptExcerptYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=true and openWorldHint=true, signaling a read operation. The description adds that inputs are validated server-side and returns a deliverable, but does not disclose additional behavioral traits like processing time or dependence on external data quality. Standard for a read tool with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with name and purpose, reference case provides concrete example. No extra fluff, but could be more structured with sections. Efficient for its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, and description only vaguely describes the return as 'structured, audited deliverable'. Missing details on possible fields, format, or behavior of the async parameter. Incomplete for a tool with nested required objects and complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17%, minimal. Description mentions three required fields (company, quarter, transcriptExcerpt) and references secFilingContext implicitly via 'case fields', but does not explain the meaning or format of each parameter beyond the reference case. Insufficient compensation for low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it produces a 'structured, audited deliverable' from earnings data and a reference case is given. The verb 'reviewer' combined with the resource implies analysis, but the purpose could be more specific about the deliverable format and audience.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus sibling tools like 'earnings_transcript_signals' or 'sec_filing_decoder'. The mention of 'FUNDRAISING' hints at context but does not provide clear when-to-use or when-not-to-use instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

earnings_transcript_signalsA
Read-only
Inspect

Earnings call transcript signal extractor for equity research analysts, catalyst-driven hedge funds, and BD teams. Parses earnings transcripts (fetched or provided) to surface:

• signals (P0/P1/P2): guidance raise/cut, miss/beat vs consensus, buyback, dividend change, new product, executive change, capex shift, M&A intent, regulatory risk, competitive threat, supply chain, hiring • kpis_mentioned: Revenue, EBITDA, EPS, FCF, Gross Margin, Operating Margin with YoY/QoQ % • guidance: raised / maintained / cut / new_initiated items extracted • q_and_a_topics: top Q&A themes detected (AI strategy, China exposure, M&A pipeline, macro, etc.) • overall_tone: bullish / neutral / bearish

Sources fetched automatically: SEC EDGAR 8-K filings, Yahoo Finance earnings news, Motley Fool transcripts. If no transcript can be retrieved from any source, returns status:'failed' with an explicit warning and empty signals — never fabricated data. Accepts transcript_text override for direct analysis. Supports multilingual transcripts (de/fr/es/zh). European tickers (SAP.DE, BMW.DE) mapped to EDGAR-compatible equivalents automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
langNoLanguage hint for the transcript. Affects mock transcript language when fetch fails.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
quarterNoFiscal quarter in format Q1-2026. Defaults to the most recent past quarter.
transcript_textNoIf provided, skips all external fetches and analyses this text directly. Minimum 100 characters.
company_or_tickerYesCompany name or ticker symbol (e.g. 'Tesla', 'TSLA', 'SAP', 'SAP.DE', 'Sanofi', 'SNY'). European tickers (SAP.DE, BMW.DE) are mapped to their ADR equivalents for EDGAR lookup.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds behavior such as automatic source fetching, failure handling (returns status:'failed' with explicit warning), multilingual support, and European ticker mapping. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points, front-loads purpose, and is informative without being overly verbose. Each sentence contributes value, though it could be slightly more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multiple outputs, source fetching, failure modes, multilingual), the description completely covers return values (signals, KPIs, guidance, Q&A, tone), failure behavior, and special features. No output schema exists, so the description adequately explains the expected outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with good parameter descriptions. The description adds marginal value by explaining European ticker mapping for company_or_ticker and transcript_text override, but does not significantly extend beyond schema for other parameters. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool extracts signals from earnings call transcripts for equity research, hedge funds, and BD teams. It lists specific outputs (signals, KPIs, guidance, Q&A topics, tone) and distinguishes itself from siblings by focusing on signal extraction from transcripts with automatic fetching, while other tools like earnings_reviewer likely have different scopes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description defines target users and use cases (e.g., analyzing earnings transcripts for signals). It explains when to use transcript_text override and what happens when no transcript is found, but does not explicitly compare to sibling tools or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

economic_indicatorA
Read-only
Inspect

Return a precise macroeconomic indicator for a country — the exact figure for a market-sizing, finance or strategy workflow. Indicators: gdp_usd, gdp_per_capita, gdp_growth, inflation, unemployment, population. Source: World Bank. When to use: an agent's analysis needs an authoritative country-level economic figure. Inputs: country (ISO-2 or ISO-3 code) and indicator name.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
countryYesCountry code, ISO-2 or ISO-3 (e.g. FR, USA)
indicatorYesMacroeconomic indicator name

Output Schema

ParametersJSON Schema
NameRequiredDescription
yearYes
valueYes
sourceYes
countryYes
indicatorYes
source_urlNo
indicator_codeNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true, covering safety and data nature. The description adds that the tool returns a precise figure from the World Bank, but does not elaborate on data freshness, historical scope, or pagination. This is acceptable but not rich beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences plus a clear list of indicators. It is concise, front-loaded with purpose, and well-structured. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (3 params, output schema present), the description covers the essential use case and inputs. It could mention the optional async parameter or error handling, but the presence of output schema and high schema coverage makes the description sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with all parameters documented. The description reiterates the parameter meanings (country code, indicator name) and lists indicators, but does not add significant new meaning beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a macroeconomic indicator, lists specific indicators, mentions source (World Bank), and specifies use cases (market-sizing, finance, strategy). It distinguishes itself from a large set of sibling tools by focusing on authoritative country-level economic figures.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states 'When to use: an agent's analysis needs an authoritative country-level economic figure,' which provides clear context. However, it does not mention when not to use or suggest alternative tools, which would be helpful given the many siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

email_domain_health_checkA
Read-onlyIdempotent
Inspect

Comprehensive email domain health check: MX routing, SPF authentication, DKIM signing, DMARC policy enforcement, DNSBL blacklist status (Spamhaus/SpamCop/Barracuda), TLS certificate validity, and WHOIS registration age. Aggregates a reputation score 0-100 and generates P0/P1/P2 deliverability signals. Accepts a domain (stripe.com) or email address (info@stripe.com). Detects role-based addresses (info@, support@, admin@, noreply@) that have higher bounce rates. Detects email provider (Google Workspace, Microsoft 365, Amazon SES, etc.). P0 signals: blacklisted / no MX / TLS expired / no SPF + DMARC none. P1 signals: SPF soft-fail / no DKIM selector / DMARC no reporting. P2 signals: role-based address / TLS expires <30d / domain age <90 days. All checks are keyless (no API keys required). Cache TTL 1h. SLA <=10s p95.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
emailNoFull email address for additional checks: format validity, role-based detection (e.g. "ceo@stripe.com").
checksNoSubset of checks to run. Defaults to all 8: ["mx","spf","dkim","dmarc","blacklist","whois","tls","reputation"]. Use a subset for faster responses (e.g. ["mx","spf","dmarc","reputation"] for quick scoring).
domainYesDomain to check (e.g. "stripe.com" or "@stripe.com"). If an email address is provided here, the domain is extracted automatically.

Output Schema

ParametersJSON Schema
NameRequiredDescription
mxYes
spfYes
tlsNo
dkimYes
dmarcYes
whoisNo
domainYes
statusYes
sourcesYes
blacklistYes
email_validNo
quality_scoreYes
reputation_scoreYes
email_is_role_basedNo
deliverability_signalsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description fully discloses behavioral traits beyond annotations: it details all 8 checks, P0/P1/P2 signal definitions, detection of role-based addresses and email providers, async behavior, caching, and SLA. Annotations only provide readOnlyHint, openWorldHint, idempotentHint; the description adds rich operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is comprehensive but well-structured: it starts with the main purpose, lists checks, explains signals, then covers parameters and operational details. It is slightly long but every sentence adds value, and it is front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of an output schema, the description is complete. It explains all inputs, outputs, behavior (async, caching, SLA), and the meaning of signals. The agent has sufficient context to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the parameter descriptions in the schema already cover the details. The tool description adds no new information about parameters beyond what is in the schema; it only reiterates. However, it does provide overall output semantics (score, signals) which aids parameter understanding indirectly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: comprehensive email domain health check including MX, SPF, DKIM, DMARC, blacklist, TLS, WHOIS, and reputation scoring. It specifies verb 'health check' and resource 'email domain', and distinguishes from siblings by focusing on email deliverability.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for checking email domain health and mentions keyless operation, caching, and SLA, but does not explicitly state when to use this tool over alternatives or provide exclusions. It lacks explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enps_autoC
Read-only
Inspect

eNPS automatisé — Gapup agent-payable C-suite expertise (CHRO). Returns a structured, audited deliverable. Reference case: BlaBlaCar — eNPS pulse mensuel · 700 FTE 8 pays · segments × tenure × manager · plays correctifs ciblés. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
contextYes
toolStackYes
segmentationYes
presenterScriptNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true, so the description's burden is low. It adds that the tool returns a deliverable and validates inputs server-side, but doesn't disclose any additional behavioral traits like processing time or output size. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (4 sentences) and front-loaded with purpose, but the use of French and a reference case slightly reduces efficiency. Every sentence adds some value, though the reference case is arguably extraneous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (7 parameters, nested objects, no output schema, 14% coverage), the description is severely incomplete. It does not describe the return value format, how to interpret the deliverable, or any post-processing steps. The agent is left with insufficient context to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% (only 'async' has a description). The description says 'send the documented case fields' but does not explain the meaning or usage of parameters like company, segmentation, context, toolStack, focus, or presenterScript. The schema structure partially guides but insufficiently.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title 'eNPS automatisé' and description mention automated eNPS analysis for CHRO/C-suite, with a reference case adding concrete context. However, the description lacks a specific verb-resource statement like 'calculates and returns eNPS analysis'; 'Returns a structured, audited deliverable' is vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the many HR-related siblings (e.g., talent_intelligence, churn_defender). It only states 'Inputs are validated server-side — send the documented case fields,' which is not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

esg_audit_multiA
Read-only
Inspect

Multi-mode ESG intelligence for ESG analysts, sustainability officers and impact investing fund managers. Aggregates live data from CDP, SBTi, Wikipedia, Yahoo Finance and web search across five modes: • company_score — ESG score 0-100 with E/S/G breakdown + heuristic rating (AAA-CCC), from CDP grade + SBTi + sector profile • controversy_check — controversies detected via web search, classified P0/P1/P2 by type (greenwashing, emissions fraud, labour, governance) • emissions — GHG Scope 1/2/3 estimates, SBTi validation flag, net-zero target year, carbon intensity per M€ revenue • esrs_readiness — CSRD gap across 12 standards (E1-E5, S1-S4, G1-G3): readiness % + gap list + CSRD deadline + effort man-days • sfdr_classification — suggested SFDR Article 6/8/9 with rationale and sustainability indicators met

Signals: P0=critical (controversy/score<40), P1=significant (score<55/SBTi missing/ESRS<50%), P2=watch. Cache 24h.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesAnalysis mode.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesCompany name, ticker, ISIN or LEI (e.g. "Microsoft", "Sanofi", "Volkswagen").
pillarNoESG pillar filter (optional, default: all).
frameworkNoESG framework filter (optional, default: all).

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
signalsYes
sourcesYes
emissionsNo
company_scoreNo
controversiesNo
quality_scoreYes
esrs_readinessNo
sfdr_classificationNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false, so the description adds value by detailing data sources (CDP, SBTi, etc.), signal severity levels (P0/P1/P2), cache duration, and per-mode behavior. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bulleted modes and a separate signals section. However, it is verbose (200+ words) and could be trimmed by removing redundant phrasing. The information density is good but not maximally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 modes, multiple data sources, signal system, async support, cache), the description covers all key aspects: what each mode returns, data sources, signal levels, and performance characteristics. No gaps are apparent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are already described. The description adds context for the 'async' parameter and lists modes in prose, but this largely repeats the schema enum. It does not significantly augment parameter understanding beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Multi-mode ESG intelligence' and enumerates five specific modes with distinct outputs (company_score, controversy_check, etc.). It differentiates itself from sibling tools by covering multiple ESG analysis types in one tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The target audience is specified (ESG analysts, sustainability officers, impact investing fund managers) and each mode's output is described, implying when to use each. However, explicit guidance on when not to use this tool versus alternatives like supplier_esg_audit or carbon_footprint_calculator is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

esrs_narrative_builderC
Read-only
Inspect

Architecte du narratif ESRS / CSRD — Gapup agent-payable C-suite expertise (SUSTAINABILITY). Returns a structured, audited deliverable. Reference case: L'Oréal France — narratif ESRS E1+E5 + S1 + G1 · CSRD reporting 2025-2026 · double-matérialité chiffrée. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
scopeYes
companyYes
contextYes
presenterScriptNo
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states it 'builds' and 'architects' a narrative, suggesting a mutating operation, but the annotation declares readOnlyHint=true, indicating a read-only operation. This contradiction reduces transparency. Additionally, it does not explain other behavioral traits like permissions or rate limits beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, but the first sentence is jargon-heavy (French/English mix) and includes unnecessary details like the reference case. It could be more concise and structured for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has complex nested parameters (company, scope, context) and no output schema. The description vaguely mentions a 'structured, audited deliverable' but does not explain return format, pagination, or how errors are handled. This is insufficient for such complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17%, meaning most parameters have no descriptions in the schema. The description does not add any meaning to the parameters; it only refers to 'documented case fields' without elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it builds ESRS/CSRD narrative for sustainability reporting and returns a structured, audited deliverable. It also provides a reference case, making the purpose specific. However, it does not explicitly differentiate from sibling tools like 'sustainability_report' or 'sustainability_reporting_pilot'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for constructing ESRS narrative and notes that inputs are validated server-side. It does not provide explicit when-to-use or when-not-to-use guidance, nor does it mention alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

event_marketingB
Read-only
Inspect

Marketing événementiel — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Pennylane (€120k/an budget événements) — 7 événements sélectionnés · coût-MQL -38% vs année précédente. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
teamSizeYes
geographyYes
objectivesYes
currentEventsYes
targetAudienceYes
annualBudgetEurYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint and openWorldHint. Description adds that inputs are validated server-side and returns a deliverable, consistent with readOnlyHint. No additional behavioral traits disclosed beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences plus a reference case – concise and front-loaded. The reference case adds valuable context but might be extraneous for agents.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no explanation of return format, and no mention of the async parameter despite its importance. The tool has nested objects and enums; description does not cover how to construct inputs. Incomplete for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (13%) and the description does not explain any parameters. It only mentions 'documented case fields' without elaboration, leaving agents without guidance for the 7 required inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it provides C-suite expertise for event marketing and returns a structured, audited deliverable. It is clear but could be more specific about the exact action (e.g., analyze, generate, evaluate). It distinguishes from many siblings but not explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implied usage for event marketing strategy, but no explicit when-to-use or when-not-to-use. No alternatives mentioned. The reference case gives context but does not guide selection vs siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

executive_comp_peer_benchmarkA
Read-onlyIdempotent
Inspect

As a Chief Human Resources Officer (CHRO), benchmark executive compensation packages against peer companies using public SEC filings and private compensation data from Equilar and Bloomberg. Inputs include executive name, title, company ticker, and peer group criteria. Outputs structured compensation metrics (base salary, bonus, equity, total compensation) with source attribution and confidence scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
peerGroupNo
fiscalYearNo
companyTickerYes
executiveNameYes
executiveTitleYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
compensationNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint. The description adds context about data sources (SEC, Equilar, Bloomberg) and output structure (metrics with attribution and confidence scores), which is useful but does not disclose potential delays or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences) and front-loaded with the primary action. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only benchmarking tool with an output schema, the description covers the essential aspects: purpose, inputs, data sources, and output types. It is complete enough for an AI agent to understand the tool's function, though it could mention the async parameter behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (17%). The description lists the key inputs (executive name, title, company ticker, peer group criteria) but does not explain the semantics of less obvious parameters like fiscalYear or peer group sub-fields. It provides partial compensation for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: benchmarking executive compensation against peer companies using specific data sources. It is specific and actionable, but does not explicitly differentiate from related sibling tools like 'comp_benchmark_geo_delta' or 'comp_plan_architect'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context ('As a CHRO, benchmark...') but provides no explicit guidance on when not to use this tool or alternatives. It lacks direct sibling differentiation or exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

financial_model_3statementA
Read-only
Inspect

Pure-compute 3-statement financial model builder (Income Statement + Balance Sheet + Cash Flow). Feed assumptions (revenue growth, COGS%, OpEx, CapEx, working capital, tax rate, depreciation, debt schedule) → receive a full 3-5 year projection with integrated DCF valuation. Supports IFRS / US_GAAP / PRC_GAAP (中国会计准则) norms with bilingual ZH+EN labels for PRC. Modes: build (full 3-statement model) | scenario_analysis (base/bull/bear ±20% growth) | sensitivity (1 KPI × 1 input, 5-point grid). No external data needed — all computed from assumptions. ICP: VC due diligence, M&A analysts, CFO SMB, startup founders pitching investors, biotech/SaaS modeling. Returns balance_check_ok per year, DCF enterprise/equity value, and coherence warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesbuild = full 3-statement model | scenario_analysis = base/bull/bear | sensitivity = 1 KPI × 1 input
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
assumptionsYesFinancial assumptions for the model
sensitivity_kpiNoKPI to observe in sensitivity mode.
sensitivity_inputNoAssumption param to vary in sensitivity mode. E.g. 'growth_rates_pct[0]' or 'cogs_pct_of_revenue'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
normsYes
statusYes
sourcesNo
warningsYes
cash_flowNo
scenariosNo
sensitivityNo
balance_sheetNo
quality_scoreYes
valuation_dcfNo
income_statementNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the tool is 'pure-compute' (read-only), requires no external data, and returns specific outputs like balance checks and coherence warnings, adding value beyond annotations which already indicate read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is fairly long but each sentence provides useful information (purpose, modes, norms, returns, ICP). It is well-structured but could be slightly more concise without losing content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, nested assumptions, output schema), the description covers all aspects: modes, accounting norms, target audience, and return values. No gaps identified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% coverage with descriptions. The description adds high-level context about norms and modes but does not provide additional meaning for individual parameters beyond what the schema already offers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it is a '3-statement financial model builder' and lists the three statements (Income Statement, Balance Sheet, Cash Flow). It clearly distinguishes from sibling tools like 'margin_doctor_finance' or 'working_capital' by focusing on comprehensive projection and DCF valuation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use (VC due diligence, M&A analysts, CFO SMB, etc.) and what inputs are needed. However, it does not explicitly compare to alternative tools or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fraud_detectorC
Read-only
Inspect

Détecteur de fraude — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: TechManu SAS — Industriel FR €32M CA, 148 FTE · 30j · 21 anomalies · €487k à risque. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
analysisPeriodDaysYes
transactionVolumesYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the tool is read-only. The description adds 'Returns a structured, audited deliverable' and mentions server-side validation, but does not disclose execution time, side effects, or authentication requirements. With annotations already covering read-only behavior, the description adds minimal value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short (3 sentences) but includes an unhelpful reference case and begins in French, which may not be appropriate for an international agent. The structure front-loads purpose but could be more concise and English-only.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, nested objects, no output schema), the description is incomplete. It does not explain how to construct inputs, what the output contains, or how the tool integrates with other tools. The reference case provides a partial example but is not systematic.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'async' parameter has a description). The description includes a reference case that hints at parameter values (company name, sector, revenue, etc.), but does not systematically explain each parameter or nested object fields. This is insufficient for a tool with 5 parameters and nested objects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The name 'fraud_detector' clearly indicates fraud detection, and the description mentions 'Détecteur de fraude' and returning a structured deliverable. However, the description is in French and includes jargon ('Gapup agent-payable C-suite expertise') that may confuse, and it does not explicitly state the tool's scope or output format. Purpose is adequate but vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. Sibling tools include similar fraud detection tools (e.g., 'affiliate_fraud_clickstream_detector', 'x402_payment_fraud_detector'), but no differentiators or usage contexts are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_business_ideasA
Read-only
Inspect

Return vetted, automation-scored business ideas from the FTG idea bank — each with an autonomy score, monetization model and conservative/median/optimistic MRR projections. When to use this tool: an agent or founder wants ranked, buildable business ideas. Input: optional category and limit.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNo
categoryNoOptional category filter

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
ideasYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds value by detailing the output structure (autonomy score, monetization model, MRR projections), which is not in annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: first states outcome, second gives usage guidance and input. Front-loaded with key information, zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values are documented. Description covers purpose, usage, and input. However, it does not mention the async parameter behavior, which is only in the schema. Still sufficient for core use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (async and category described). The description simply repeats 'optional category and limit' without adding details beyond the schema. Limit parameter lacks description in both schema and description, so no extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('return') and resource ('vetted, automation-scored business ideas') with clear output details (autonomy score, MRR projections). It distinguishes from siblings like ftg_business_plan and ftg_market_gap by focusing on idea discovery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'an agent or founder wants ranked, buildable business ideas.' Provides context but does not mention when not to use or explicitly name alternatives, though the sibling list implies differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_business_planA
Read-only
Inspect

Return the business plan for a market-gap opportunity — direct-trade or local-production, with CAPEX, OPEX, ROI, payback period, automation level and the full plan. Cache-first: returns the stored plan when available, otherwise reports that generation is required (the FTG platform produces plans on demand). When to use this tool: an agent has an opportunity_id (from ftg_market_gap) and needs the investable plan. Input: an opportunity_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
opportunity_idYesOpportunity id obtained from ftg_market_gap

Output Schema

ParametersJSON Schema
NameRequiredDescription
plansNo
statusYes
messageNo
plan_countNo
opportunity_idYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate read-only and open-world behavior. Description adds cache-first behavior and explains that if not cached, it reports generation required. No contradictions; provides useful behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise and well-structured: first sentence states output and contents, second explains cache behavior, third gives usage guidance. No wasted words, information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description adequately covers what the tool does, its prerequisites, and behavior. It is complete for a tool that returns a business plan and fits within the suite of ftg tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by clarifying that opportunity_id comes from ftg_market_gap, which is crucial for correct usage. Async parameter is already well-described in schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a business plan for a market-gap opportunity, listing included elements (CAPEX, OPEX, ROI, payback period, automation level) and distinguishing it from siblings by specifying it uses opportunity_id from ftg_market_gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: when an agent has an opportunity_id and needs the investable plan. Also mentions cache-first behavior. Could be more explicit about when not to use, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_country_regulationsA
Read-only
Inspect

Return import, trade and production regulations for a country — category, title, summary and source. When to use this tool: an agent checks regulatory or compliance requirements before trading or producing in a market. Input: a country, with an optional category.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNo
countryYesCountry ISO code or name
categoryNoOptional regulation category filter

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
regulationsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark it as read-only and open-world. The description restates that it returns data but does not add extra behavioral context like caching, rate limits, or data freshness. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences plus a usage line and an input line. Every sentence serves a distinct purpose—describing output, usage context, and input parameters. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists (per context signals) and the tool is a straightforward read-only lookup, the description provides enough context for an agent to understand its role. It could mention what happens when no regulations are found, but the output schema likely handles that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explicitly says 'Input: a country, with an optional category,' which adds value by making the optionality clear beyond the schema's property descriptions. However, it does not cover optional 'async' or 'limit' parameters, though these are common patterns and documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns import, trade, and production regulations for a country, listing specific fields (category, title, summary, source). This distinguishes it from the many sibling tools that deal with other aspects of trade, compliance, or country data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes a 'When to use' statement that gives context (regulatory/compliance checks before trading/producing), but it does not explicitly mention when not to use it or point to alternative tools for related but distinct tasks (e.g., sanctions screening, trade finance eligibility).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_country_studyA
Read-only
Inspect

Return the in-depth FTG country study — multi-part structured analysis of a country's trade and production landscape. When to use this tool: an agent needs deep country context before a sourcing, export or investment decision. Input: a country.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
countryYesCountry ISO code or name

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
partsYes
countryYes
part_countYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. Description adds that it's a structured analysis but doesn't disclose traits like speed or async behavior beyond the schema. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences effectively conveying purpose and usage. No wasted words, front-loaded with key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given annotations, output schema existence, and full param coverage, the description is sufficient. It could mention the async parameter context but is not necessary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and describes both parameters well. Description only adds 'Input: a country,' which is redundant. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns an in-depth multi-part structured analysis of a country's trade and production landscape, and specifies the use case. It does not explicitly differentiate from sibling tools like ftg_country_regulations, but the purpose is distinct enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: when deep country context is needed for sourcing, export, or investment decisions. Does not mention when not to use or alternatives, but provides clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_investor_directoryA
Read-only
Inspect

Return investors from the FTG directory — VC, PE and impact funds with type, firm, website, ticket-size range, sectors and stages of interest. When to use this tool: an agent builds a fundraising shortlist. Input: optional country and limit.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNo
countryNoOptional country ISO code or name

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
investorsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true. Description adds that it returns investors with specified fields and optional inputs, but does not disclose additional behavioral traits such as pagination, rate limits, or data freshness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is two sentences, front-loaded with purpose, then usage and input. No redundant or unnecessary information. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has output schema, so return values are covered. Description lists returned fields (type, firm, etc.) and notes optional inputs. Lacks details on default limit, handling of multiple countries, or pagination, but is largely adequate for a directory lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% (async and country described; limit only min/max). Description mentions 'optional country and limit' but does not add meaning beyond schema for country and provides no details on limit behavior (e.g., default value). Async parameter is not mentioned in description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Return investors from the FTG directory' with specific data fields (type, firm, etc.). It distinguishes from siblings like 'investor_list' by specifying the source (FTG directory) and the scope of data returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'When to use this tool: an agent builds a fundraising shortlist.' Provides clear context but does not mention when not to use or suggest alternatives like 'investor_shortlist' or other investor tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_market_gapA
Read-only
Inspect

Return the import/production market-gap opportunities for a country — commodities where local demand outpaces local supply. Each opportunity carries the gap value (USD/year), the gap volume (tonnes/year), a 0-100 opportunity score and the potential margin. When to use this tool: an agent needs to know what a country structurally under-produces or over-imports, for trade sourcing, import/export or local-production investment decisions. Input: a country (ISO-2 code or name).

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum opportunities to return (default 20)
countryYesCountry ISO-2 code (e.g. 'SN', 'KE') or name (e.g. 'Senegal')

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
countryYes
opportunitiesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. Description adds output details (gap value, volume, score, margin) but no further behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise, front-loaded with purpose, each sentence adds value. No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Has output schema, so return values covered. Description mentions key output fields. Could mention behavior when no gaps found or limit clamping.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%; description adds minimal value beyond schema, just restating country input format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns market-gap opportunities for a country, with specific output fields. It distinguishes itself from siblings by focusing on import/production gaps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use (trade sourcing, import/export decisions) but lacks when-not-to-use or alternatives, despite many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_opportunity_scoutA
Read-only
Inspect

Rank the best countries for a given commodity — where the market gap, opportunity score and potential margin are highest. Cross-country scouting. When to use this tool: an agent has a commodity and needs to know WHERE to sell, export to or set up local production. Input: a commodity name.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum countries to return (default 20)
commodityYesCommodity name (e.g. 'rice', 'soybean', 'poultry')

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
countYes
commodityYes
countriesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. The description adds no behavioral details beyond purpose and usage, missing opportunities to mention data sources or limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with purpose and usage. No redundant or tangential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return value details are not needed. Description covers purpose and usage adequately, though a brief comparison to similar ftg tools could enhance completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. Description only mentions 'Input: a commodity name' and does not add meaning beyond the schema's parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool ranks countries by market gap, opportunity score, and potential margin for a given commodity, distinguishing it from sibling tools like ftg_market_gap or ftg_production_economics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: when an agent has a commodity and needs to know where to sell, export, or set up production. Does not mention alternatives or when not to use, but context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_production_economicsA
Read-only
Inspect

Return production cost benchmarks (CAPEX/OPEX per unit, value ranges, scenarios, quality tiers) and agronomic yields (t/ha, cycles per year) for a commodity. When to use this tool: an agent sizes the economics of producing a commodity. Input: a commodity, with an optional country.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNo
countryNoOptional country ISO code or name
commodityYesCommodity name or slug

Output Schema

ParametersJSON Schema
NameRequiredDescription
yieldsYes
commodityYes
cost_benchmarksYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, indicating a safe, read-only operation with variable results. The description adds value by specifying the type of data returned (costs, yields) but does not disclose other behavioral traits like rate limits, pagination, or async behavior (though async is in the schema). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences plus a usage note, with essential information front-loaded. Every sentence adds value, and there is no redundancy or fluff. It achieves maximum efficiency for its purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only retrieval tool with an output schema (indicated by context signals), the description covers the main input (commodity), optional parameter (country), and usage context. It does not detail the output schema, but that is acceptable given its existence. The tool's complexity is moderate, and the description, combined with annotations and schema, provides sufficient guidance for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with 'limit' lacking a description in the schema. The description provides some additional context by noting 'Input: a commodity, with an optional country.' This partially compensates but does not explain all parameters (e.g., async, limit). Baseline 3 is appropriate given the high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns production cost benchmarks and agronomic yields for a commodity, using a specific verb 'Return' and specifying the resource. While it does not explicitly distinguish itself from siblings like 'ftg_production_methods' or 'ftg_market_gap', the purpose is unambiguous and appropriate for an economics-sizing task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit usage context: 'When to use this tool: an agent sizes the economics of producing a commodity.' This provides clear guidance on the intended use case. However, it does not mention when not to use it or suggest alternative tools, which would strengthen the dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_production_methodsA
Read-only
Inspect

Return the production methods for a commodity — each with a description, ordered process steps, pros/cons and a popularity rank. Methods are commodity-canonical: one curated set per commodity, reusable across every country. When to use this tool: an agent evaluates HOW a commodity is produced or processed, compares techniques, or builds a production plan. Input: a commodity slug or name.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
commodityYesCommodity slug or name (e.g. 'rice', 'tomato', 'cashew')

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
methodsYes
commodityYes
method_countYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds context about return content (steps, pros/cons, rank) but does not disclose any behavioral traits beyond what annotations provide. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (4 sentences) and front-loaded with the core action. Every sentence adds value: output, nature, use cases, input. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description provides sufficient context: purpose, input, output nature, and usage. It also includes the nuance of commodity-canonical methods. Slightly more detail about the return structure could elevate it, but it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for both parameters. The description reiterates the input as 'commodity slug or name' but adds no new meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Return the production methods for a commodity' and specifies the output details (description, steps, pros/cons, rank). It distinguishes from siblings like ftg_production_economics by focusing on the 'how' of production and mentions commodity-canonical nature, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage scenarios: 'evaluates HOW a commodity is produced or processed, compares techniques, or builds a production plan.' It does not list when not to use or alternatives, but the context is clear enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_seller_catalogA
Read-only
Inspect

Return seller catalogues registered on FTG — exporters and producers with their commodity, monthly capacity, certifications and target export markets. When to use this tool: an agent builds a supplier or sourcing shortlist. Input: optional seller country and commodity.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNo
countryNoOptional seller country ISO code or name
commodityNoOptional commodity filter

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
sellersYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavioral context beyond the annotations (readOnlyHint, openWorldHint) by detailing the returned data fields and optional filters. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise with two sentences plus a usage hint, all front-loaded. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and the presence of an output schema, the description covers all necessary context: what the tool returns, when to use it, and optional inputs. No gaps in essential information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, meeting the baseline. The description mentions optional country and commodity filters but does not add significant new meaning beyond what the schema already provides for those parameters. It does not clarify the 'limit' or 'async' parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns seller catalogues with specific attributes like commodity, capacity, certifications, and export markets. It distinguishes the tool's function from related siblings like 'ftg_sourcing_buyers' by specifying the exact data returned, though it does not explicitly differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'when an agent builds a supplier or sourcing shortlist.' This provides clear context for use, though it does not mention when not to use or alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ftg_sourcing_buyersA
Read-only
Inspect

Return verified local buyers in a country — companies sourcing a given commodity, with buyer type, city, website, annual volume range and certification requirements. When to use this tool: an agent builds a sourcing or export shortlist, or needs real B2B demand contacts in a market. Input: a country and an optional commodity filter.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum buyers to return (default 20)
countryYesCountry ISO-2 code or name
commodityNoOptional commodity slug to filter buyers by

Output Schema

ParametersJSON Schema
NameRequiredDescription
buyersYes
countryYes
commodityNo
buyer_countYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, which establish a read-only, potentially non-exhaustive data source. The description aligns with these (returning 'verified local buyers') and adds the nuance of async execution. No contradiction, and the description provides context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (two sentences plus a usage note), front-loaded with the core purpose, and free of redundant or irrelevant information. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (4 parameters, required country only, output schema present), the description covers the key aspects: input, output, usage context, and async behavior. It lacks explicit error handling or pagination instructions, but the limit parameter and async mechanism are documented in the schema. Overall complete for a lookup tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description only reiterates 'Input: a country and an optional commodity filter' without adding new details or clarifying format beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return verified local buyers'), the specific resource (country and commodity), and the output fields (buyer type, city, website, annual volume range, certification requirements). This distinguishes it from sibling tools like ftg_investor_directory or ftg_seller_catalog, which target different roles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage context: 'When to use this tool: an agent builds a sourcing or export shortlist, or needs real B2B demand contacts in a market.' It does not explicitly name alternative tools for sellers or investors, but the sibling list and the description's focus on 'buyers' imply differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

funding_hunterC
Read-only
Inspect

Chasseur de financements — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: PME deeptech cleantech FR €8M CA — top 30 dispositifs BPI+France2030+EU+VC. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
projectYes
financialsYes
preferencesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true. Description adds no further behavioral context beyond stating it returns a deliverable, which is consistent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is short and front-loaded, but lacks substantive content. Efficiency is not beneficial when key information is missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested parameters and lack of output schema, the description is insufficient. It does not explain return value structure or interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only async parameter described). The description provides no detail on any parameter, failing to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a 'funding hunter' returning a structured deliverable, with a specific reference case. However, it does not differentiate from similar sibling tools like 'capital_strategy' or 'investor_list'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives, nor when not to use it. Only implies sending documented case fields.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fx_rateA
Read-only
Inspect

Get the current or historical foreign-exchange rate for any currency pair — the exact exchange rate, FX rate or conversion rate an agent needs to convert a currency amount or feed a finance, trading, invoicing or pricing workflow. Covers EUR/USD, USD/JPY, GBP/EUR and every ISO-4217 currency pair. Returns the latest spot rate, or a historical rate by date. Use when a workflow needs a precise live or past currency exchange rate, or to convert money between two currencies. Source: European Central Bank reference rates via Frankfurter. Inputs: from/to ISO-4217 currency codes, optional date (YYYY-MM-DD).

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesQuote currency, ISO-4217 (e.g. USD)
dateNoOptional YYYY-MM-DD for a historical rate
fromYesBase currency, ISO-4217 (e.g. EUR)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toYes
fromYes
rateYes
as_ofYes
sourceYes
source_urlNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint as true. The description adds context: returns spot rate or historical rate, source is ECB via Frankfurter, and date format. It does not contradict annotations. It adds value beyond annotations by explaining the data source and output specifics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single but well-structured paragraph. It front-loads the main purpose, then adds scope, use case, source, and inputs. Every sentence adds value, and there is no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (so return values are covered) and simple parameters, the description is quite complete. It covers purpose, usage, source, and input details. It could be more precise about ECB rate limitations (e.g., base currency EUR), but overall it is informative enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description reinforces 'from/to ISO-4217 currency codes' and 'optional date (YYYY-MM-DD)' but does not add meaning beyond what the schema provides. The async parameter is not mentioned in the description, missing an opportunity to clarify its use.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool gets current or historical foreign-exchange rates for any ISO-4217 currency pair. It specifies the verb (get) and resource (rate) and provides examples. While it does not explicitly differentiate from sibling tools like supply_chain_fx_exposure_dashboard, the purpose is specific and distinct enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use: 'Use when a workflow needs a precise live or past currency exchange rate, or to convert money between two currencies.' This gives clear context. It does not provide when-not-to-use or alternatives, but given the sibling list has no direct replacement, this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

geographic_expansionC
Read-only
Inspect

Expansion géographique — Gapup agent-payable C-suite expertise (CSO). Returns a structured, audited deliverable. Reference case: Gapup Hub — Expansion 4 marchés (DE/UK/ES/NL) · €1.8M budget · ARR cible €3.2M Y2. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
productYes
financialsNo
constraintsNo
targetMarketsYes
preferredEntryModeNo
expansionHorizonMonthsNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. The description adds context about server-side validation and a structured deliverable, which is consistent and mildly informative, but does not disclose detailed behavior like cost, auth, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and front-loaded with the core purpose. The reference case adds context but is not essential. Could be slightly more concise without the example, but overall it's efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (8 parameters, nested objects, no output schema), the description is too sparse. It does not elaborate on the deliverable format, return structure, or how to use the parameters effectively, leaving significant gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 13%, yet the description provides no parameter-level details beyond 'send the documented case fields'. Most parameters (e.g., financials, constraints) are left unexplained, forcing the agent to rely on the sparse schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly indicates this tool provides a structured, audited deliverable for geographic expansion, with a reference case. However, it does not explicitly differentiate from similar tools like market_entry_strategist, which weakens clarity slightly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. The instruction 'send the documented case fields' is vague and does not help the agent decide when this tool is appropriate over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

geo_logistics_intelA
Read-only
Inspect

Geospatial logistics intelligence for supply chain, maritime and transport agents. Four modes: (1) geocode_batch — resolve up to 50 addresses to lat/lon with confidence scores (OSM Nominatim + Open-Meteo fallback, 1 req/s rate-limit respected); (2) routing — road/cycling/walking route with distance_km, duration_seconds and ETA ISO timestamp between two addresses or lat/lon points (OSRM public, keyless, global); (3) port_congestion — congestion status for any UN/LOCODE port (e.g. NLRTM, SGSIN, CNSHA) with waiting vessel count, severity (low/medium/high/extreme) and average wait hours; (4) ship_tracking — AIS position, speed, course, destination and ETA for a vessel by its 9-digit MMSI. No API key required for geocode/routing/port. Optional env: AIS_STREAM_API_KEY for live ship data (otherwise MarineTraffic scrape best-effort). SLA: <=25s p95. Cache: 24h geocoding / 1h routing / 30min port / 5min ship. Quality score 0-100. Status: final/partial/failed.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNorouting only: destination address or 'lat,lon'
fromNorouting only: origin address or 'lat,lon'
modeYes'geocode_batch': address -> lat/lon. 'routing': route + ETA. 'port_congestion': UN/LOCODE port state. 'ship_tracking': vessel by MMSI
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesPrimary input: address for geocode/routing, UN/LOCODE (e.g. NLRTM) for port_congestion, 9-digit MMSI for ship_tracking
addressesNogeocode_batch only: up to 50 addresses (overrides query if provided)
mode_transportNorouting only: transport mode. Default: driving

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
routingNo
sourcesYes
geocode_batchNo
quality_scoreYes
ship_trackingNo
port_congestionNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and destructiveHint=false, which the description aligns with. The description adds substantial behavioral context beyond annotations: rate limits (1 req/s), SLA (≤25s p95), caching durations (24h geocoding, 1h routing, 30min port, 5min ship), fallback mechanisms (Open-Meteo), and best-effort behavior for ship tracking without API key. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized as a single paragraph listing modes numerically. Every sentence adds value, though bullet points or subsection headers could improve readability. It is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 modes), complete annotations, and full schema coverage, the description covers all necessary aspects: input formats, output details (lat/lon, route parameters, congestion metrics, AIS data), performance characteristics, cache durations, and fallback behavior. It is comprehensive and leaves no significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds significant meaning: it explains the role of each parameter in context (e.g., 'addresses overrides query for geocode_batch'), provides examples (NLRTM, lat,lon), and clarifies constraints (max 50 addresses). This goes beyond the schema's property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Geospatial logistics intelligence for supply chain, maritime and transport agents.' It explicitly lists four modes (geocode_batch, routing, port_congestion, ship_tracking) with specific verbs and resources, distinguishing it from siblings which are unrelated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear when-to-use guidance for each mode (e.g., 'resolve up to 50 addresses to lat/lon', 'route between two points', 'congestion status for any UN/LOCODE port', 'AIS position for a vessel by MMSI'). It also specifies prerequisites (no API key needed for most, optional for ship tracking) and constraints (rate limits, SLA, cache durations). However, it does not explicitly state when not to use this tool or compare to alternative tools, leaving room for improvement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

global_salary_inflation_adjusterA
Read-onlyIdempotent
Inspect

Adjusts salary benchmarks for local inflation using OECD, IMF, and World Bank data. Designed for CHROs to normalize compensation across regions with accurate inflation adjustments. Inputs include country codes, base salary, and reference year. Outputs inflation-adjusted salary with data sources and warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
baseSalaryYes
targetYearNo
countryCodeYes
referenceYearYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
targetYearNo
countryCodeNo
inflationRateNo
referenceYearNo
adjustedSalaryNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations by specifying data sources (OECD, IMF, World Bank) and output content (data sources and warnings). Annotations already mark the tool as readOnlyHint, idempotentHint, and openWorldHint, so the description does not contradict them instead, it enriches the behavioral understanding. Score 4 because it provides useful additional context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with only three sentences, each adding value. It front-loads the core action and target users, then briefly lists inputs and outputs. No unnecessary words, making it easy to read and understand quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (not shown but referenced), the description's high-level mention of 'inflation-adjusted salary with data sources and warnings' is sufficient. Annotations cover safety and idempotency. However, the description does not explain when to use the 'async' parameter, which is a minor gap for completeness. Overall, the description is fairly complete for a read-only, idempotent tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only the 'async' parameter has a description). The description mentions 'country codes, base salary, and reference year' as inputs, adding some meaning beyond the schema for these required parameters. However, it does not explain 'targetYear' or the 'async' parameter's purpose (though async is described in schema). The description partially compensates for low schema coverage but not fully, so a score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool adjusts salary benchmarks for local inflation using OECD, IMF, and World Bank data, and is designed for CHROs to normalize compensation across regions. The verb 'adjusts' combined with the specific resource 'salary benchmarks' makes the purpose clear. However, it does not explicitly differentiate from sibling tools like comp_benchmark_geo_delta, so it loses a point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context about intended users (CHROs) and the goal (normalize compensation across regions), which implies when to use. However, it offers no guidance on when not to use or explicit alternatives. Sibling tools exist but are not mentioned, so the usage guidance is adequate but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gl_reconcilerC
Read-only
Inspect

GL Reconciler — Réconciliation grand livre — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Answers: Identify the root causes of the GL breaks in 's ledger for — cluster them and rank by materiality. · For Q close: which accounts have unreconciled items over €? Provide a sign-off routing and resolution plan. · Run an automated GL reconciliation for — AR/AP/intercompany entries — flag open items, suggest journal entries. · What are the top 5 systemic control weaknesses causing recurring GL breaks at ? Recommend preventive controls. · Generate a month-end close reconciliation report for — breaks by account type, aging analysis, sign-off assignments. Reference case: Acme SaaS Q4 2026 — 47 breaks GL, €1.4M variance non postée. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
entityYes
ledgerContextYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side and that a structured deliverable is returned, which is consistent with read-only behavior. However, it does not disclose further behavioral traits like rate limits, required permissions, or how results can be polled (despite the async parameter).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long and unstructured, mixing French and English, marketing language, and a reference case ('Acme SaaS Q4 2026...'). It could be condensed into a clear single sentence about the tool's function, followed by parameter explanations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the general purpose and gives concrete examples of what the tool can answer, which is helpful given the lack of output schema. However, it does not specify the output structure or how to correctly fill the nested parameters (e.g., 'entity' fields), leaving gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 4 parameters with only 25% description coverage (async has a description). The description does not explain the 'entity', 'focus', or 'ledgerContext' parameters or how they map to the example queries. With such low schema coverage, the description should compensate but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool performs GL reconciliation and returns a structured deliverable. It lists example user queries, such as identifying root causes of GL breaks and generating month-end close reports, making the purpose clear. However, it does not explicitly state that it is a read-only analysis tool, relying on annotations for that.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides example queries but gives no guidance on when to use this tool versus its many siblings (e.g., financial_model_3statement, audit_pre_flight). No explicit when-to-use or when-not-to-use instructions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gov_procurement_multiA
Read-only
Inspect

Aggregate public procurement tenders (calls for tender / appels d'offres) from multiple government sources simultaneously: TED Europa v3 (27 EU countries, keyless API), BOAMP France (opendatasoft, keyless), UK Contracts Finder (OCDS standard, keyless), SAM.gov United States (requires SAM_GOV_API_KEY env var), and bund.de Germany (HTML scraping, partial). Returns structured tender records with buyer authority, EU CPV sector code, estimated contract value converted to EUR via live FX rates, submission deadlines, and direct notice URLs. Use when: a B2G agent needs to find government contract opportunities matching keywords across multiple jurisdictions; building a pipeline of public tenders for bid/no-bid qualification; monitoring a domain by CPV code; market sizing public sector spend. Key inputs: query (keywords), countries (ISO-2 array), cpv_codes (EU standard codes, e.g. 72000000=IT services, 45000000=construction, 79000000=business services), min_value_eur (filter), published_after (ISO date, defaults to 30 days ago). SLA: <=25s p95 (all sources fetched in parallel, 8s budget per source). Optional env var SAM_GOV_API_KEY enables US federal tenders (free key at api.sam.gov). Quality score: 25 pts if TED EU retrieved, 15 pts per other source retrieved (max 60), 10 pts if >= 10 tenders returned, 5 pts if aggregates computed. Status: failed < 30 / partial 30-59 / final >= 60.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesKeywords to search for tenders (e.g. "cybersecurity audit", "construction", "consulting AI")
countriesNoCountries to search. Defaults to ["EU","US","FR","UK","DE"]. Use "EU" for all 27 EU member states via TED Europa.
cpv_codesNoEU Common Procurement Vocabulary codes (e.g. ['72000000'] for IT services, ['45000000'] for construction). Optional.
min_value_eurNoMinimum contract value in EUR. Tenders below this are excluded. Optional.
published_afterNoISO date YYYY-MM-DD. Only return tenders published after this date. Defaults to 30 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
queryYes
statusYes
sourcesYes
tendersYes
by_sourceYes
by_countryYes
quality_scoreYes
countries_searchedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnly, not destructive), description adds behavioral details: parallel fetch, SLA 25s, quality scoring, optional env var, and status levels. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is comprehensive but logically structured with lists and key sections. It is slightly verbose but each sentence adds value. Good front-loading of purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters, multiple sources, and complex behavior (parallel fetch, quality scoring), the description covers return structure, SLA, error states, env var requirement, and limits, making it complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and descriptions are present. The tool description adds examples (e.g., CPV codes like 72000000), default behaviors (published_after defaults to 30 days), and clarifies ISO-2 format, enhancing schema info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it aggregates public procurement tenders from multiple named government sources, with details on what is returned. It clearly differentiates from sibling tools by its multi-jurisdiction scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use when' section provides clear use cases (B2G needs, pipeline building, CPV monitoring, market sizing) and lists key inputs. It does not explicitly exclude alternative tools but context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

growth_path_architectC
Read-only
Inspect

Architecte de croissance — Gapup agent-payable C-suite expertise (CSO). Returns a structured, audited deliverable. Reference case: Pennylane (€30M ARR) — 3 voies de croissance · Mix recommandé : Organique + Geo EU · ARR cible €120M en 36 mois. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
constraintsYes
growthTargetYes
currentDriversYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint and openWorldHint, which are consistent with the description's claim of returning a deliverable. The description adds that inputs are validated server-side, but no additional behavioral context (e.g., cost, latency, data sources) beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the purpose. The reference case adds some context but could be considered extraneous. Overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the inputs (5 parameters, deep nesting) and lack of output schema, the description does not provide enough context on what the deliverable contains or how to interpret results. An agent would lack understanding of the tool's output and return structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 20%, and the description does not explain any parameter meanings beyond 'send the documented case fields.' The complex nested inputs (company, growthTarget, etc.) are left to the schema alone, which is insufficient for an agent to know how to fill them properly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is for growth architecture, targeting C-suite expertise, and returns a structured deliverable. It gives a reference case (Pennylane). However, it does not differentiate from many similar strategy planning siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. The description only says to send documented case fields, but does not indicate prerequisites, when not to use, or comparison to siblings like market_entry_strategist or strategic_options_analyzer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hallucination_confidence_meterA
Read-onlyIdempotent
Inspect

Evaluates the likelihood of hallucination in LLM responses by comparing against HuggingFace model confidence scores. Designed for risk assessment personas to quantify response reliability. Accepts text snippets or model outputs, returns confidence metrics and potential hallucination warnings. Cross-references with top-performing models from the HuggingFace leaderboard.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe LLM-generated text to evaluate for hallucination risk
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
model_idNoOptional specific HuggingFace model ID to use for evaluation
thresholdNoConfidence threshold below which hallucination warnings are triggered

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
confidence_scoresNo
hallucination_likelihoodNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations: it accepts text snippets or model outputs, returns confidence metrics and potential hallucination warnings, and cross-references with top HuggingFace models. It does not contradict any annotations (readOnlyHint, openWorldHint, idempotentHint). No mention of rate limits or auth needs, but these are not expected from annotations either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: three sentences that efficiently convey the purpose, target users, inputs, and outputs. Every sentence adds value with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 parameters, output schema exists, annotations provided), the description covers all necessary information: what it does, for whom, what it takes, and what it returns. No gaps are apparent, and it is complete for an agent to decide whether to use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for all 4 parameters, so the baseline is 3. The description does not add new meaning beyond what the schema already provides for each parameter (text, async, model_id, threshold). It mentions 'text snippets or model outputs' which aligns with the text parameter but adds no new semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool evaluates hallucination likelihood in LLM responses by comparing against HuggingFace model confidence scores. It specifies the target audience (risk assessment personas) and the output (confidence metrics and warnings). This distinguishes it well from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions it is designed for risk assessment personas, giving a clear context of use. However, it does not explicitly state when not to use this tool or provide alternatives among the many sibling tools, such as bias_amplification_tracker or jailbreak_attempt_detector.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

historical_price_seriesA
Read-onlyIdempotent
Inspect

Fetch historical OHLCV price series for any ticker: stocks (AAPL, SAP.DE, 7203.T), ETFs, indices, commodities (GC=F for gold) or cryptocurrencies (BTC-USD). Returns a full date-indexed series of open/high/low/close/volume plus pre-computed statistics: total return, annualised return (CAGR), annualised volatility, max drawdown and Sharpe estimate (rf=4%). Automatically detects crypto tickers (→ CoinGecko) vs traditional assets (→ Yahoo Finance primary, Stooq fallback). Adjusts for dividends and splits when adjusted=true (default). Use cases: backtesting, factor analysis, performance attribution, charting, financial modelling. Sources: Yahoo Finance, CoinGecko, Stooq. All keyless. Optional env: AICI_RESEARCH_PROXY_URL for Bright Data routing (lifts Yahoo 429), TWELVE_DATA_API_KEY for higher Twelve Data quota.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
periodNoLook-back period. Default: 1y.
tickerYesYahoo Finance ticker symbol. Examples: AAPL (US stock), SAP.DE (Frankfurt), 7203.T (Tokyo), BTC-USD (Bitcoin), GC=F (gold futures), ^GSPC (S&P 500).
metricsNoSubset of fields to include (informational — all fields always returned).
adjustedNoAdjust close prices for dividends and splits. Default: true.
intervalNoBar interval. Default: 1d (daily).

Output Schema

ParametersJSON Schema
NameRequiredDescription
statsYes
periodYes
seriesYes
statusYes
tickerYes
sourcesYes
currencyYes
intervalYes
data_pointsYes
quality_scoreYes
splits_detectedNo
resolved_exchangeNo
dividends_detectedNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description adds significant context beyond annotations: auto-detects crypto vs traditional assets, adjusts for dividends/splits, mentions fallback sources, keyless access, optional proxy for rate limits, and async behavior. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is a single paragraph but packs essential information efficiently. It is front-loaded with purpose. Slightly long but every sentence adds value; could be broken into bullet points but not required.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, description covers sources, use cases, rate limit handling, and ticker formats. It is comprehensive for a complex tool with 6 parameters and diverse use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description enriches parameter meanings: ticker examples, adjusted defaults, metrics being informational, async explanation. Adds value beyond schema especially for ticker and async.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Fetch historical OHLCV price series for any ticker' with specific examples (stocks, ETFs, indices, commodities, cryptocurrencies). It distinguishes the tool's broad scope, though it does not explicitly differentiate from sibling tools. However, the specificity and completeness of purpose earn a top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description lists explicit use cases ('backtesting, factor analysis, performance attribution, charting, financial modelling') and mentions auto-detection of ticker types. It does not provide when-not-to-use guidance, but the context is clear enough for most scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hr_benefits_esg_alignerA
Read-onlyIdempotent
Inspect

Asynchronous tool for Chief Human Resources Officers (CHROs) to align employee benefits packages with ESG (Environmental, Social, Governance) goals. Uses Eurostat HR data, MSCI ESG ratings, and Sustainalytics metrics to generate actionable recommendations. Inputs include company location, industry, and current benefits structure. Outputs ESG-aligned benefits adjustments with sustainability impact scores. Requires async:true to avoid timeout errors.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
esgFocusNoPrimary ESG pillars to prioritize
industryCodeYesNACE or ISIC industry classification code
companyLocationYesISO 2-letter country code of company headquarters
currentBenefitsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
recommendationsNo
overallESGAlignmentScoreNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations: it is async, uses specific data sources (Eurostat, MSCI, Sustainalytics), and outputs recommendations with scores. This adds value as the annotations only indicate readOnly, openWorld, and idempotent. There is no contradiction between the description and annotations (readOnlyHint is plausible for a recommendation generator).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact three-sentence paragraph that front-loads purpose, then describes inputs/outputs and a key behavior (async). Every sentence contributes value without redundancy. Slightly more structure could improve scannability, but it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (not shown but indicated), the description adequately covers inputs, data sources, and async behavior. It provides enough context for an AI agent to understand when and how to invoke it, though it could mention the output format briefly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, high enough for a baseline of 3. The description mentions the parameters (company location, industry, current benefits) but adds little detail about their meaning beyond what the schema already provides. It does explain the async parameter's purpose, which is helpful. Overall, marginal added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: aligning employee benefits with ESG goals for CHROs. It specifies the verb 'align', the resource 'benefits packages', and the target audience. This distinguishes it from sibling tools like procurement_okr_esg_aligner which focus on procurement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions the tool is asynchronous and requires async:true to avoid timeouts, and that it uses specific data sources. However, it does not explicitly state when to use this tool versus alternatives, nor does it provide conditions for when not to use it. The guidance is limited to the async behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

incident_response_evidence_collectorA
Read-onlyIdempotent
Inspect

As a CTO, gather forensic evidence (logs, network flows, MITRE TTPs) from public breach reports and threat intelligence sources to support incident response post-mortems. Inputs include incident identifiers, date ranges, or MITRE technique IDs. Outputs structured evidence with attack patterns, indicators of compromise, and source references. — pass async:true REQUIRED to avoid x402 timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
date_rangeNo
incident_idYesUnique identifier for the incident (e.g., CVE, GitHub Advisory ID)
mitre_technique_idsNoList of MITRE ATT&CK technique IDs (e.g., T1059)
include_network_flowsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
timelineNo
warningsNo
indicatorsNo
incident_idNo
network_flowsNo
attack_patternsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds important behavioral context beyond annotations, specifically the requirement to pass async:true to avoid x402 timeout. The annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the description supplements these with operational constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, front-loading the purpose and then specifying inputs and outputs. The inclusion of 'As a CTO' is slightly unnecessary but does not detract much. The async note is placed at the end, effectively highlighting a critical usage requirement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and annotations are present, the description covers the essential aspects: purpose, input types, output structure, and a key behavioral note (async requirement). It does not mention error handling or pagination, but for a read-only evidence collector, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description adds limited semantic value by mentioning input types (incident identifiers, date ranges, MITRE technique IDs) and the important async parameter. It does not explain all parameters in detail, so it does not significantly elevate the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the tool's purpose: gathering forensic evidence from public breach reports and threat intelligence to support incident response post-mortems. It mentions specific evidence types (logs, network flows, MITRE TTPs) and outputs structured evidence, which distinguishes it from sibling tools that focus on other aspects like vulnerability scanning or compliance audits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (incident response post-mortems) and the required inputs (incident identifiers, date ranges, or MITRE technique IDs). However, it does not explicitly state when not to use it or mention alternative tools among the many siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

india_market_dataA
Read-only
Inspect

Indian capital market intelligence for the IN diaspora (30M+) and investors. Covers NSE, BSE, and MCA corporate registry across four modes:

• company — full company profile: name, CIN, exchange, NSE/BSE tickers, industry, incorporation date, paid-up capital, registered office, status, directors • market_quote — real-time quote: price (INR), change%, volume, market cap, P/E ratio. Sources: Yahoo Finance (primary), BSE API, NSE API (proxy-gated) • sector_overview — Nifty/Sensex sector snapshot: top 5 companies by market cap. Supported sectors: it, banking, pharma, energy, auto, fmcg, realestate, metals, telecom, consumer • mca_filing — Ministry of Corporate Affairs filings. Requires CIN for direct lookup.

Input formats accepted: • NSE ticker (e.g. 'RELIANCE', 'TCS.NS') • BSE 6-digit code (e.g. '500325' for Reliance) • CIN 21-char (e.g. 'L17110MH1973PLC019786') • Company name EN (e.g. 'Reliance Industries', 'Tata Consultancy') • Sector keyword (e.g. 'IT services', 'banking', 'pharma')

ENV: AICI_RESEARCH_PROXY_URL with country-in routing unlocks NSE direct API and MCA.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesAnalysis mode.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesNSE/BSE ticker, CIN (21 chars), company name (EN), or sector keyword.
exchangeNoExchange filter. Default: all.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
queryYes
statusYes
companyNo
sourcesYes
mca_filingsNo
market_quoteNo
quality_scoreYes
sector_overviewNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable context: it mentions real-time data, sources (Yahoo Finance, BSE API), async mode for slow queries, and environment variable requirements. This goes beyond annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with bullet points and clear sections. Every sentence adds value, from purpose to modes to input formats. It is front-loaded with the main action and avoids waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 modes, multiple input formats, API sources), the description covers all necessary aspects: modes, inputs, environment setup. The presence of an output schema (not shown) further reduces burden. No gaps are evident.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds significant meaning: it explains each mode's output, details input formats (ticker, CIN, name, sector), and provides examples. The async and exchange parameters are also clarified, making the schema more actionable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides 'Indian capital market intelligence' and lists four specific modes (company, market_quote, sector_overview, mca_filing) with concrete details. This distinguishes it from siblings like china_market_data and gives a clear purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the scope (NSE, BSE, MCA) and modes, but does not explicitly contrast with other tools or state when not to use it. The context is clear enough for an agent to decide, but lacks explicit alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

industry_classifier_naics_sicC
Read-only
Inspect

Classificateur d'industrie NAICS/SIC/NACE — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Answers: What is the NAICS code for a company that does ? · Give me NAICS + SIC + NACE classification for this company description. · Which industry sector (GICS) does this company belong to for equity analysis? · What HS code applies to products manufactured by this company? · For EU procurement compliance, what NACE Rev. 2 code applies to this company? · Classify this business into NAICS + SIC + ISIC + GICS + NACE + HS with hierarchy and confidence. · I need to segment my ICP list by NAICS 4-digit subsector — classify these company descriptions. Reference case: Helios Cold Chain EU — Freight forwarding maritime réfrigéré · . Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
company_urlNo
company_nameNo
company_descriptionYes
focus_classificationsNo
primary_revenue_sourceNo
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. The description adds minimal behavioral context beyond mentioning server-side validation and that it returns a structured deliverable. No significant additional disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat verbose with many example questions, but the core purpose is front-loaded. It could be more concise without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 6 parameters and no output schema, the description is incomplete. It does not explain return format, error handling, or how to interpret results beyond 'structured, audited deliverable'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (17%). The description does not provide meaningful explanations of parameters beyond listing example queries. Key parameters like focus_classifications are not elaborated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it classifies companies into NAICS/SIC/NACE and returns a structured deliverable, with example queries. However, it does not explicitly differentiate from sibling tools, which are numerous but this one seems unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for any company description to get industry codes but lacks explicit when-to-use or when-not-to-use guidance, and no alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

infra_blueprint_designerB
Read-only
Inspect

Architecte infra cloud — Gapup agent-payable C-suite expertise (CTO). Returns a structured, audited deliverable. Answers: Design a cloud infrastructure blueprint for a app with expected traffic and requirements. · What is the recommended AWS vs GCP vs Azure architecture for a SaaS multi-tenant app with EU data residency and SOC2? · How should I architect my cloud infra to stay under €5k/month with GDPR compliance and a junior DevOps team? · What cloud services do I need for a with load — compute, DB, cache, CDN, observability? · Give me an end-to-end cloud architecture with scaling plan, security baseline, and IaC tool recommendation. Reference case: Spinora fintech B2B SaaS — saas-multi-tenant · medium load (1k-100k req/d) · eu-west · . Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
team_sizeNo
expected_loadYes
workload_typeYes
business_contextNo
cloud_preferenceNo
region_preferenceYes
budget_monthly_eurNo
compliance_requiredNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds valuable context: it is an 'Architecte infra cloud' that provides CTO-level expertise, mentions the ability to use async mode ('returns a job_id immediately'), and notes server-side validation. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is excessively verbose, containing a long block of example questions and mixed French/English. It lacks structure and conciseness; the key information could be conveyed in fewer sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the tool's complexity (9 parameters, no output schema), the description fails to explain the return format, limitations, or expected behavior beyond stating it returns a 'structured, audited deliverable'. More completeness is needed for an agent to use it reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11%, meaning most parameters lack explanations. The description does not systematically describe each parameter beyond the example queries. For a tool with 9 parameters and low schema coverage, more explicit parameter semantics are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Architecte infra cloud — ... Returns a structured, audited deliverable.' It includes specific example queries that demonstrate its function (e.g., 'Design a cloud infrastructure blueprint for a <workload_type> app'). The title and examples distinguish it from sibling tools, none of which focus on cloud architecture design.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implicit usage guidance through example questions, indicating when to use the tool (e.g., for cloud architecture design with specific constraints). However, it does not explicitly state when not to use the tool or mention alternative tools, leaving the agent to infer from context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

insurance_coverage_analyzerB
Read-only
Inspect

Analyseur de couvertures d'assurance — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: Gapup Hub — 3 polices · €24k prime · Score 58/100 · 3 gaps critiques · RFP template. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
arrEurYes
sectorYes
objectivesYes
companyNameYes
riskProfileYes
jurisdictionYes
employeeCountYes
currentPoliciesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. Description adds return type (structured, audited deliverable) and validation behavior. Could elaborate on output structure and error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is short and front-loaded with purpose. The reference case adds specificity but is somewhat cryptic. Efficient use of words, no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite complexity (nested objects, 9 params, no output schema), description lacks details on output format, scoring, or gap identification. The reference case hints but does not fully equip an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (11%). Description only vaguely refers to 'documented case fields' and a reference example, but fails to explain any of the 9 parameters in detail. Compensation is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it analyzes insurance coverage and returns a structured, audited deliverable. The reference case example reinforces purpose. However, no explicit differentiation from siblings, but the name and French description make it stand out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a reference case example implying typical usage scenarios. Mentions server-side validation. But lacks explicit when-not-to-use or alternatives among numerous sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

interest_rateA
Read-only
Inspect

Return a precise reference interest rate — the exact figure an agent injects into a treasury, lending, valuation or trading model. Available rates: fed_funds, sofr, us_10y, us_2y, us_3m, ecb_main, euribor_3m. Source: FRED (Federal Reserve Bank of St. Louis). When to use: an agent's computation needs a current benchmark rate as a precise input.

ParametersJSON Schema
NameRequiredDescriptionDefault
rateYesReference rate name
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.

Output Schema

ParametersJSON Schema
NameRequiredDescription
rateYes
unitYes
as_ofYes
valueYes
sourceYes
series_idNo
source_urlNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description doesn't need to duplicate that. It adds context about the data source (FRED) and that the rate is precise, but doesn't disclose additional behavioral traits like update frequency or caching.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences plus a list, front-loading the action and purpose. Every sentence contributes value with zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with an output schema, the description provides the key context: exact rates available, source, and use case. It could mention that rates are current/latest, but overall it's adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description lists the available rates which matches the enum in schema, and mentions 'precise reference interest rate' but doesn't add distinct semantic value beyond what the schema provides for each parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Return a precise reference interest rate' and lists exactly which rates are available (fed_funds, sofr, etc.) and the source (FRED). This distinguishes it from sibling tools like fx_rate or economic_indicator which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'When to use: an agent's computation needs a current benchmark rate as a precise input.' This gives clear context, though it does not mention when not to use or alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

internal_communicationC
Read-only
Inspect

Communication interne — Gapup agent-payable C-suite expertise (CHRO). Returns a structured, audited deliverable. Reference case: Cas démo — Communication interne. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
contextYes
audienceSegmentsYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint=true and openWorldHint=true, which already indicate a safe read operation. The description adds that inputs are validated server-side and that it returns a 'structured, audited deliverable', but does not elaborate on behavioral traits such as authentication needs, rate limits, or what happens to existing data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the tool's purpose, but it includes a vague reference case and a note on validation. It could be more concise and informative, sacrificing clarity for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex input schema (nested objects, multiple fields) and the lack of an output schema, the description is insufficient. It does not clarify the expected output structure or how inputs map to the deliverable, leaving significant gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%. The description mentions 'documented case fields' but does not explain any specific parameters or their roles. With low coverage, the description fails to compensate, leaving parameter meanings unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Communication interne — Gapup agent-payable C-suite expertise (CHRO)' and says it 'Returns a structured, audited deliverable.' This clearly indicates the tool generates a report for internal communication. However, it lacks an explicit verb like 'generate' or 'create', and its purpose is not sharply differentiated from similar HR-related sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It only instructs to 'send the documented case fields', but does not specify contexts, prerequisites, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

investor_listB
Read-only
Inspect

Liste d'investisseurs + warm intros — Gapup agent-payable C-suite expertise (FUNDRAISING). Returns a structured, audited deliverable. Reference case: Agicap Série D — 25 VCs matchés · Tier A: Balderton/Accel/Partech · Warm intro path chaque investisseur. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
roundYes
companyYes
existingInvestorsNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true, so the tool is likely safe and not destructive. The description adds that it returns a 'structured, audited deliverable' and mentions async capability (via job_id), but does not contradict annotations. It provides moderate additional context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a short paragraph with a clear front-loaded purpose. However, it mixes French and English, and includes redundant phrases (e.g., 'Inputs are validated server-side'). It could be more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested objects, no output schema), the description lacks details on the return format or structure of the 'audited deliverable'. It does not fully prepare the agent to interpret results. More completeness is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only 25% of parameters have schema descriptions (async has a description). The description does not explain the 'round', 'company', or 'existingInvestors' fields beyond stating to send 'documented case fields'. With low schema coverage, the description should compensate, but it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it provides a list of investors with warm intros for fundraising, and references a specific use case (Agicap Series D). However, it does not explicitly differentiate this tool from siblings like 'investor_shortlist' or 'funding_hunter'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for fundraising (C-suite expertise) and mentions server-side validation, but does not provide explicit guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

investor_shortlistC
Read-only
Inspect

Shortlist d'investisseurs ciblés — Gapup agent-payable C-suite expertise (FUNDRAISING). Returns a structured, audited deliverable. Reference case: Aleph AI — Series B €30M · 60 investisseurs EU/US matchés par stage/thèse · fit score + warm intro path + first message angle. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
roundYes
companyYes
preferencesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. The description adds that the tool returns a 'structured, audited deliverable' but lacks details on processing time, external calls, or cost implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description includes promotional language and an example, making it longer than necessary. It is moderately concise but could be more direct.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested schema and lack of output schema, the description should explain the output structure more concretely. It mentions fit score, intro path, and message angle but doesn't specify the deliverable's format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 20% and the description does not elaborate on parameter meanings beyond 'send the documented case fields'. Nested objects in the schema are left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it creates a shortlist of targeted investors for fundraising, with a reference case and mention of a structured deliverable. However, it does not explicitly differentiate from sibling tools like 'investor_list'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this vs. alternatives. The description mentions server-side validation but doesn't specify when to use the tool or when not to, leaving the agent to infer usage from context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ip_contract_clause_extractorA
Read-onlyIdempotent
Inspect

For CHRO use: analyzes employment contract text to identify and extract IP-related clauses such as invention assignment, confidentiality, non-compete, and patent rights. Returns structured data with clause types, risk levels, and relevant legal context. Ideal for contract review workflows, compliance checks, and IP protection strategy. Sources: USPTO PatFT and EPO Espacenet public datasets. Keywords: employment contract, IP clause, invention assignment, confidentiality agreement, non-compete, patent rights, CHRO tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
contractTextYesFull text of the employment contract to analyze
jurisdictionNoCountry/state jurisdiction for legal context (e.g., 'US-CA', 'DE')
includeContextNoWhether to include legal context for each clause

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
clausesYes
sourcesNo
summaryYes
warningsYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true. Description adds that it returns structured data with clause types, risk levels, and legal context, and mentions external data sources (USPTO, EPO). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences with front-loaded purpose, output description, use cases, and sources. Slightly verbose in the keywords sentence (repeats terms from first sentence), but otherwise well-structured and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description adequately covers tool behavior, input requirements, use cases, and data sources. Does not miss any critical aspects for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all 4 parameters. Description does not add new semantics beyond schema, but reinforces the purpose (e.g., 'contractText' is the full text). Baseline 3 is appropriate for high-coverage schema with minimal added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool analyzes employment contract text to extract IP-related clauses (invention assignment, confidentiality, non-compete, patent rights). Specifies target user (CHRO) and differentiates from sibling tools like legal_clause_extractor by focusing on IP clauses in employment context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides context for use: 'For CHRO use', 'Ideal for contract review workflows, compliance checks, and IP protection strategy.' Does not explicitly mention when not to use or name alternatives, but the context is clear enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ip_employee_invention_trackerA
Read-onlyIdempotent
Inspect

For CHROs: tracks employee patent filings and flags unassigned inventions. Input employee name or ID to retrieve their patent applications from USPTO and WIPO databases. Returns list of inventions with assignment status, filing dates, and potential ownership gaps. Useful for IP audits, inventor onboarding, and compliance checks. Keywords: patents, IP ownership, employee inventions, USPTO, WIPO.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
endDateNoFilter patents filed before this date (YYYY-MM-DD)
startDateNoFilter patents filed after this date (YYYY-MM-DD)
employeeIdNoInternal employee ID (optional if name provided)
companyNameYesExact legal name of company for assignment check
employeeNameYesFull name of employee to track (e.g., 'John Doe')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
patentsYes
sourcesNo
warningsYes
employeeIdNo
companyNameYes
employeeNameYes
totalPatentsYes
unassignedPatentsYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint. The description adds behavioral context beyond these, such as the output includes 'assignment status, filing dates, and potential ownership gaps', and the databases queried. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences with no extraneous text. It front-loads the main purpose and uses keywords efficiently. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the richness of annotations (readOnly, openWorld, idempotent) and the presence of an output schema, the description provides adequate high-level context about inputs, behavioral queries, and return format. It could mention the date filtering parameters, but overall it's complete enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters well. The description mentions 'employee name or ID', matching the schema, but does not add additional meaning beyond the schema for the date filters or other parameters. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('tracks', 'flags', 'retrieve') and clearly identifies the resource ('employee patent filings', 'USPTO and WIPO databases'). It distinguishes the tool's employee-specific scope from broader sibling tools like patent_landscape.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description targets a specific user ('For CHROs') and lists concrete use cases ('IP audits, inventor onboarding, compliance checks'). However, it does not explicitly state when not to use this tool or mention alternatives among the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ip_protection_pilotC
Read-only
Inspect

Pilote de protection IP — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: Carbios SA — Deeptech FR recyclage PET enzymatique · 14 brevets EP/US/FR · 5 concurrents · licensing €2-8M potentiel. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
competitorsYes
targetMarketsYes
patentPortfolioSummaryYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true. The description adds that it returns a deliverable and validates inputs server-side, but does not disclose whether results are immediate or any rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short and front-loaded with the purpose. The reference case example adds length but provides useful context, though it could be trimmed without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters with nested objects, no output schema, and no description of return format or behavior, the description is insufficient. It lacks detail on expected output, parameter constraints, and how this tool fits into a workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only 'async' documented). The description mentions 'send the documented case fields' but does not explain the meaning of parameters like company, patentPortfolioSummary, or competitors. The reference case provides some minimal context but is insufficient for the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it's an IP protection pilot that returns a structured, audited deliverable, with an example case. However, it does not explicitly differentiate from sibling IP tools like patent_landscape or patent_ownership_audit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives, nor any exclusions or prerequisites. Only states that inputs are validated server-side, providing no usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jailbreak_attempt_detectorA
Read-onlyIdempotent
Inspect

Detects potential LLM jailbreak attempts by analyzing user input against NIST AI Risk Management Framework adversarial patterns. Designed for persona risk assessment, this tool evaluates text for common jailbreak techniques such as prompt injection, role-playing, or obfuscation. Inputs include the user message and optional context, returning a risk assessment with confidence scores and pattern matches. Ideal for real-time moderation in chat applications or API gateways.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
contextNoOptional conversation context for better pattern matching
messageYesUser input text to analyze for jailbreak attempts
thresholdNoConfidence threshold for flagging attempts

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
riskScoreNoConfidence score of jailbreak attempt
patternsMatchedNoList of detected adversarial patterns
isJailbreakAttemptNoWhether the input exceeds the risk threshold
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds beyond annotations by stating the tool returns a risk assessment with confidence scores and pattern matches, and its design for persona risk assessment. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three focused sentences, front-loaded with the core purpose, and each sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema and moderate complexity (4 params, 1 required), the description adequately covers purpose, inputs, output type, and use cases, though could mention async behavior or threshold details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of parameters with descriptions; the description reiterates 'user message and optional context' without adding new meaning or constraints beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool detects jailbreak attempts using NIST AI RMF patterns, lists specific techniques (prompt injection, role-playing, obfuscation), and identifies use cases like real-time moderation, making it distinct from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions ideal use cases (chat applications, API gateways) but does not explicitly differentiate from similar tools like adversarial_input_stress_tester or safety_guardrail_breach_analyzer, nor specify when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_postings_intelligenceA
Read-only
Inspect

Agrégation d'offres d'emploi publiques pour inférer les tendances de recrutement. Trois modes : (1) company_hiring — analyse des postings d'une société : volume, fonctions (engineering/sales/marketing/ops/finance/hr), seniorité, géographie, croissance vs période précédente, signaux stratégiques inférés ; (2) role_market — volume marché global pour un rôle (open positions estimate, top employeurs, compétences demandées, médiane seniorité) ; (3) competitor_hiring_comparison — comparaison multi-sociétés (total postings, growth%, focus areas). Sources : Adzuna (ADZUNA_APP_ID/KEY env), RemoteOK (keyless), Himalayas (keyless), baseline statique 40 top employeurs. Usages : due diligence VC, intelligence compétitive, benchmarks RH, signaux pivots stratégiques. Cache 6h. SLA ≤15s.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesMode d'analyse : 'company_hiring' | 'role_market' | 'competitor_hiring_comparison'
roleNoIntitulé de poste à analyser (pour role_market, ex. 'data scientist', 'compliance officer')
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyNoNom de la société (pour company_hiring ou comme 1er concurrent)
locationNoPays ou ville (ex. 'France', 'United States', 'London')
competitorsNoListe de sociétés à comparer (pour competitor_hiring_comparison, min 2)
period_daysNoFenêtre d'analyse en jours (défaut 30)

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
sourcesYes
role_marketNo
quality_scoreYes
company_hiringNo
competitor_comparisonNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, openWorldHint) are complemented by behavioral details: cache 6h, SLA ≤15s, and async parameter (though not described in description, it's in schema). No contradiction. The description adds value beyond structured fields by explaining data freshness and performance expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed and well-structured, using bullet points for modes. While concise for the amount of information, it could be slightly trimmed (e.g., removing 'Usages' redundancy). However, it remains efficient and easy to parse for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, 3 modes, external data sources), the description covers purpose, modes, sources, cache, SLA, and intended use cases. An output schema exists, so return values are not required. It is comprehensive enough for correct tool selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions, providing a baseline of 3. The description enriches parameter understanding by explaining each mode's purpose and linking parameters to modes (e.g., 'company' for company_hiring, 'competitors' for competitor_hiring_comparison). It adds context beyond the schema's basic field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: aggregating public job postings to infer recruitment trends. It specifies three distinct modes (company_hiring, role_market, competitor_hiring_comparison) and each mode's function, effectively distinguishing it from a large set of sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists explicit use cases (due diligence VC, competitive intelligence, HR benchmarks, strategic pivot signals) and data sources (Adzuna, RemoteOK, Himalayas, static baseline). It does not explicitly state when not to use or name alternative sibling tools, but the provided modes and context give sufficient guidance for appropriate invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_resultA
Read-onlyIdempotent
Inspect

Poll the result of any tool called with async:true. Returns status=pending while running, status=completed with the full result once done, status=failed on error, or status=not_found if the job_id is unknown or expired (TTL 24h).

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by an async tool call

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, idempotent, non-destructive. The description adds value by detailing polling behavior, possible statuses (pending, completed, failed, not_found), and the 24h TTL, providing transparency beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with action and enumeration of statuses. No redundant words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description need not detail return values. It covers the essential behavior, statuses, and TTL, making it complete for a polling tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The parameter job_id is fully described in schema ('The job_id returned by an async tool call'). The description references it but adds no further meaning. With 100% schema coverage, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool polls results of async tool calls, specifying the action ('poll') and resource ('result of any tool called with async:true'). It distinguishes from sibling tools like ai_governance_full_report_result by being generic.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use ('Poll the result of any tool called with async:true'), but does not explicitly exclude alternative scenarios or compare with specific result tools. However, the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kalshi_marketsAInspect

Query live Kalshi prediction markets (CFTC-regulated US exchange). Returns question, implied probability (0-1, derived from the yes bid/ask mid), volume, open interest, close time and URL. Optional free-text filter on the question.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum markets (default 20)
queryNoFree-text filter on the market question
statusNoMarket status (default open)
includeRawNoInclude Kalshi's original fields (default false)
includeUnpricedNoAlso return markets with no live bid/ask (default false — they carry no information)
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the implied probability is derived from the yes bid/ask mid and that the exchange is CFTC-regulated. However, it does not mention pagination, rate limits, error behavior, or the meaning of 'unpriced' markets—leaving some behavioral aspects unexplored.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary action, and every sentence contributes useful information: what the tool does and what it returns. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a moderately complex schema (6 optional parameters) and no output schema, so the description must explain return values—which it does by listing the fields. It does not mention the status filter, limit, async behavior, or raw mode, but those are covered in the schema. The core purpose and return structure are adequately conveyed, though additional details on output shape (array vs. object) would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all six parameters. The description adds minimal semantics beyond the schema, only mentioning the free-text filter. It does not compensate with extra context for parameters like includeRaw or includeUnpriced, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool queries live Kalshi prediction markets, distinguishing it from sibling tools like Polymarket. It lists specific return fields (question, implied probability, volume, open interest, close time, URL) and mentions the optional free-text filter, providing a precise verb+resource definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context that this tool is for Kalshi markets, differentiating it from Polymarket or general prediction market search. However, it does not explicitly mention when not to use it or name alternatives, so it stops short of a 5. The context is clear, with no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

knowledge_base_autoC
Read-only
Inspect

Base de connaissance automatique — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: Klarna — knowledge base auto · Slack+Notion+Drive · 12 articles seed + structure 8 catégories. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
sourcesYes
topPainPointsYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and openWorldHint=true, which the description does not contradict. However, the description adds little beyond stating that inputs are validated server-side. It does not elaborate on behavioral traits such as authentication, rate limits, or what happens to the deliverable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short at a few sentences, but it includes a reference case that may not be essential. The structure is front-loaded with the tool's purpose, but the French phrases reduce clarity. Could be more concise without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of nested objects, no output schema, and 5 parameters, the description is incomplete. It does not describe the output format, how the knowledge base is structured, or how to use the async parameter effectively. The agent lacks crucial context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only async described). The description does not explain the meaning of key parameters like company, sources, topPainPoints, or focus beyond saying 'send the documented case fields'. This is insufficient for an agent to understand how to populate the nested objects correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it creates an automated knowledge base for C-suite expertise and returns a structured audited deliverable, but it does not clearly differentiate from many sibling tools that also generate reports or governance artifacts. The French phrasing and reference case add some context but lack specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The reference case provides one example but does not explain when this tool is appropriate or when to choose different tools. No when-not-to-use or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kyc_screenerC
Read-only
Inspect

Screening KYC / AML / Sanctions — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: Q4 2026 onboarding — 8 entités (UBO chain LLC + SPV offshore), sanctions/PEP/adverse media. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
entitiesYes
riskAppetiteYesstandard
screeningScopeYes
onboardingPacketYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. The description adds that it returns a structured, audited deliverable and that inputs are validated server-side, but does not cover additional behaviors like sync/async defaults or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences but contains unnecessary jargon and a long example. It could be more concise and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given nested objects and no output schema, the description lacks details on the deliverable structure, async behavior, polling mechanism (job_result), and result interpretation. Incomplete for a complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 20% schema coverage, the description should compensate by explaining parameters, but it only vaguely mentions 'send the documented case fields'. It adds little meaning beyond schema field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool screens KYC/AML/Sanctions and provides an example reference case, making the purpose specific. However, jargon like 'Gapup agent-payable C-suite expertise (RISK)' may reduce clarity for some agents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus sibling tools such as kyc_screener_batch or sanctions_screener_multi. The description implies single-case screening but does not differentiate or mention alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kyc_screener_batchA
Read-only
Inspect

Async batch variant of kyc_screener. Accepts 1-100 names and returns immediately (<300ms) with a job_id. The screening runs in the background (up to 10 parallel KYC calls). Poll the result with kyc_screener_batch_result(job_id) after the eta_seconds hint. Each entry can specify name, type (person/company/any), and an optional birthdate hint. Use for bulk client onboarding, UBO list screening, or periodic AML refresh batches. Async tool — register a webhook via webhooks_manage(register, url, [job.completed]) to receive callbacks instead of polling. Faster + lighter.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
namesYesList of entities to screen (1-100). Each entry requires at minimum a name.

Output Schema

ParametersJSON Schema
NameRequiredDescription
job_idYesUnique job identifier — pass to kyc_screener_batch_result
statusYes
batch_sizeYesNumber of names queued for screening
eta_secondsYesEstimated seconds until result is ready
submitted_atYesISO-8601 submission timestamp
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the description adds value by explaining the async run, job_id, eta_seconds hint, and webhook option. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph but well-organized: purpose, behavior, use cases, alternatives. Efficient with no wasted words, though could benefit from more structured formatting.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex async tool with an output schema (not shown), the description covers input, behavior, output (job_id), polling, and webhook alternative. Sufficient for an agent to decide usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline 3. The description adds meaning by explaining names accept 1-100, type enum values, and birthdate for disambiguation, plus the async parameter's effect. This exceeds baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it's an async batch variant of kyc_screener, accepting 1-100 names and returning a job_id. It distinguishes from the synchronous single by specifying batch size and async behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context for use (bulk onboarding, UBO list screening, AML refresh) and mentions an alternative (webhook callback). It lacks explicit exclusion for when to use the non-batch variant but implies it through context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kyc_screener_batch_resultA
Read-onlyIdempotent
Inspect

Poll the result of a kyc_screener_batch job. Returns status=pending while running, status=completed with the full array of KYC results once done, status=failed on error, or status=not_found if the job_id is unknown or expired (TTL 24h). Call this after the eta_seconds hint returned by kyc_screener_batch.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by kyc_screener_batch (prefix: kycb_)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the read-only nature is clear. The description adds value by disclosing the TTL (24h expiry) and explaining each status outcome, which goes beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single well-structured paragraph that lists statuses and gives usage guidance. Every sentence adds value, no wasted words, and it is front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's polling nature with multiple status outcomes and a TTL, the description covers all necessary behavioral aspects. It also references the parent batch tool, making the polling flow complete. The output schema existence further reduces burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with the job_id parameter fully described (type, required, prefix format). The description adds polling behavior context but does not enrich parameter semantics beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it polls the result of a KYC screener batch job, enumerating possible statuses (pending, completed, failed, not_found) and TTL. It ties directly to its sibling tool kyc_screener_batch, distinguishing its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises 'Call this after the eta_seconds hint returned by kyc_screener_batch', providing clear context on when to use it. No exclusions or alternatives mentioned, but the guidance is specific.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

labor_law_alert_geoA
Read-onlyIdempotent
Inspect

Provides CHROs with daily alerts on new labor law changes by jurisdiction (state/country). Inputs include jurisdiction (ISO country/state code) and optional date range. Outputs structured legislative updates with summaries, effective dates, and source links. Useful for compliance monitoring, risk assessment, and policy adjustments. Keywords: labor law, compliance, legislation, jurisdiction, CHRO, HR policy.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
sinceNoOptional start date for changes (YYYY-MM-DD). Defaults to 7 days ago.
untilNoOptional end date for changes (YYYY-MM-DD). Defaults to today.
jurisdictionYesISO 3166-1 alpha-2 country code or ISO 3166-2 state/province code (e.g., 'US-CA', 'FR')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
changesYes
sourcesYes
warningsYes
last_updatedNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, indicating safe reads. The description adds that the tool returns structured legislative updates with summaries, effective dates, and source links, which provides useful behavioral context beyond the annotations. No contradictions are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences plus keywords, front-loaded with the main purpose. No redundant information; every sentence adds value. It is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema, the description adequately explains the output content (structured updates with summaries, effective dates, source links). Annotations cover safety and idempotency. The tool is relatively simple, and the description provides sufficient context for agent decision-making.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with all parameters described in the input schema. The description mentions jurisdiction and optional date range, aligning with the schema but not adding significant new meaning. The jurisdiction format is already specified in the schema. Baseline is 3, and the description provides marginal additional clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: providing daily alerts on new labor law changes by jurisdiction to CHROs. It specifies the input (jurisdiction and optional date range) and output (structured legislative updates with summaries, effective dates, source links). This distinguishes it from sibling tools like 'compliance_monitor' and 'legal_clause_extractor' by focusing on jurisdiction-specific legislative changes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists use cases such as compliance monitoring, risk assessment, and policy adjustments, providing context for when to use. However, it does not explicitly mention when not to use or compare to alternatives among the many sibling tools. The keywords help infer usage but lack direct exclusions or comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ld_architectC
Read-only
Inspect

Architecte formation & développement — Gapup agent-payable C-suite expertise (CHRO). Returns a structured, audited deliverable. Reference case: Pennylane (180 FTE) — Catalogue 8 formations · 3 parcours individuels · ROI €480k · Payback 7 mois. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
teamYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
budgetYes
companyYes
learningNeedsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so the description does not need to restate safety. It adds that the tool returns a structured audited deliverable and inputs are validated server-side. However, it does not detail what the deliverable contains (e.g., format, schema) or any additional behavioral traits like async behavior beyond what the async parameter already conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise, consisting of two sentences plus a reference case. It is front-loaded with the purpose. The mix of French and English may reduce clarity slightly, but overall it is efficient. The reference case adds length but also value for context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has complex nested input, no output schema, and many siblings. The description provides a reference case and target audience, but it does not explain the output format, how to use the result, prerequisites, or what qualifies as a valid input beyond schema constraints. This leaves significant gaps for an agent to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (20%), with only the async parameter having a description. The description says 'send the documented case fields' but does not explain what each field (company, team, learningNeeds, budget) means or how to fill them correctly. The agent is left to infer from parameter names alone, which is insufficient for proper invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it is an architect for training & development, returning a structured audited deliverable, with a specific reference case. The verb is implicit ('returns'), but the resource is clear. It does not explicitly differentiate from siblings like 'lnd_ai_skill_forecast' or 'lnd_skill_taxonomy_builder', but the title and reference to C-suite expertise and ROI suggest a strategic planning role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for CHRO-level strategic planning and mentions a reference case, but it does not explicitly state when to use this tool vs alternatives. There is no guidance on when not to use it or comparison with sibling tools, leaving the agent without clear selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lead_magnetsC
Read-only
Inspect

Aimants à leads — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Spendesk — Guide trésorerie startup SaaS B2B FR/EU (2024). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
icpYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
brandYes
leadMagnetYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and openWorldHint=true. The description says 'Returns a structured, audited deliverable,' which may imply generation but not necessarily a write operation. No conflicts with annotations. However, the description fails to disclose behavioral traits such as typical execution time, authentication needs, or whether it modifies data. With annotations covering safety, the description adds minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at three sentences, but one sentence is in French, and the structure is disjointed. It front-loads the tool name but lacks logical flow. Could be more efficient and clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 4 complex nested parameters and no output schema. The description does not specify what the deliverable contains, how it is structured, or what to expect. With no output schema and high parameter complexity, the description is incomplete for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only 'async' parameter described). The description does not explain any of the required parameters (icp, brand, leadMagnet) or their structure. Given low coverage (<50%), the description should compensate but does not, leaving agents to infer from nested schemas without semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title and description indicate it's a lead magnet tool that returns a structured deliverable, referencing a specific case. However, the purpose is somewhat vague: it mixes French and English and does not clearly state what a lead magnet is in this context or how it differs from other content generation tools among many siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. It mentions inputs are validated server-side, but does not specify prerequisites or conditions for use. The sibling list is long, but the description does not help differentiate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lgpd_data_subject_rights_automatorB
Read-onlyIdempotent
Inspect

Automates LGPD Data Subject Access Requests (DSARs) for legal teams, handling Brazil-specific data retention, erasure, and access workflows. Accepts user identifiers, request type (access/rectification/deletion), and optional scope filters. Returns structured response with compliance status, warnings, and source references to Brazilian LGPD and CNIL decisions.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
scopeNoOptional list of data categories to limit the request
urgencyNoPriority level for processing
requestTypeYesType of LGPD request
userIdentifierYesCPF, email, or other unique identifier for the data subject

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
dataCategoriesNo
erasureDeadlineNo
complianceStatusNo
retentionPeriodDaysNo
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, idempotentHint=true, but the description mentions handling erasure and deletion, which contradicts read-only semantics. Additionally, it does not disclose the async behavior indicated by the async parameter. This contradiction and lack of detail result in poor transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no wasted words: first sentence explains purpose, second lists inputs, third describes outputs. It is front-loaded and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description fails to mention the async behavior, which is crucial for proper usage. It also contradicts annotations by implying mutation (deletion) despite readOnlyHint. This incompleteness undermines its usefulness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description summarizes inputs (user identifiers, request type, scope filters) but adds no new detail beyond what the schema already provides. It does not explain the async or urgency parameters. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it automates LGPD DSARs for legal teams, specifying Brazil-specific data retention, erasure, and access workflows. It uses a specific verb 'Automates' and identifies the resource, distinguishing it from other privacy tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies it is for Brazil-specific LGPD workflows, but does not explicitly state when to use this tool over alternatives or when not to use it. It lacks guidance on exclusions or comparison with other DSAR tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lnd_ai_skill_forecastA
Read-onlyIdempotent
Inspect

Forecasts AI skill demand trends for CHROs by analyzing patent filings (USPTO PatFT) and job postings (BLS API). Returns 12-month skill demand projections with confidence scores, helping HR leaders prioritize workforce upskilling. Inputs: target AI skills (e.g., 'machine learning', 'NLP'), geographic focus (US state/country), and forecast horizon. Outputs include skill growth rates, patent filing trends, and job posting volumes. Keywords: AI workforce planning, skill gap analysis, talent strategy, patent trends, labor market data.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
regionYesGeographic focus (US state code or 'US' for national, e.g., 'CA', 'US')
skillsYesList of AI-related skills to forecast (e.g., ['machine learning', 'computer vision'])
horizon_monthsNoForecast horizon in months (3-24, default 12)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
forecastNo
metadataNo
warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds valuable context: it uses patent and job data, returns 12-month projections with confidence scores, and lists specific outputs. This goes beyond what annotations provide, though no detailed behavioral caveats are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single three-sentence paragraph that front-loads the purpose and audience. It is reasonably concise but includes some redundancy (e.g., 'Inputs:' and 'Outputs:' lists reiterate schema). Still, it is well-structured and avoids fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a complete input schema (100% coverage) and an output schema, the description provides sufficient context: audience, data sources, output types, and use case. It fully explains what the tool does and what the output contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions. The description mentions the parameters generically ('target AI skills', 'geographic focus') but does not add deeper semantics beyond what the schema already provides, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'forecasts AI skill demand trends for CHROs' using specific data sources (USPTO PatFT and BLS API). It distinguishes itself from sibling tools like 'job_postings_intelligence' and 'patent_landscape' by combining both datasets and focusing on AI skills, making its purpose unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the target audience (CHROs) and use case (prioritizing workforce upskilling). However, it does not provide guidance on when not to use this tool or mention alternatives among the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lnd_skill_taxonomy_builderA
Read-onlyIdempotent
Inspect

Generates a dynamic skill taxonomy for CHROs by cross-referencing patent filings (USPTO), job postings (BLS), and learning & development data (OECD). Inputs include industry codes, job roles, or skill clusters; outputs structured skill hierarchies with demand trends and competency gaps. Essential for workforce transformation, talent pipeline optimization, and future-proofing organizational capabilities. — pass async:true REQUIRED to avoid x402 timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
jobRoleNoTarget job role or occupation (e.g., 'Data Scientist')
industryYesNAICS industry code or sector name (e.g., '541511' for IT services)
timeRangeNoTime range for trend analysis
skillClusterNoOptional skill cluster to focus taxonomy (e.g., 'AI/ML')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
skillTaxonomyNo
industryTrendsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, openWorldHint, and idempotentHint. The description adds behavioral context by detailing data sources and the mandatory async flag. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences and a key usage note. It front-loads the primary action. Some marketing language is present but not excessive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and annotations, the description covers the tool's purpose, data sources, input/output types, and the critical async requirement. It does not explain polling but the output schema likely provides return structure context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with adequate descriptions for each parameter. The description summarizes input types but adds no new parameter-level details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates a dynamic skill taxonomy for CHROs using specific data sources (USPTO, BLS, OECD), and specifies inputs (industry codes, job roles, skill clusters) and outputs (structured skill hierarchies with demand trends and competency gaps). This differentiates it from sibling tools like lnd_ai_skill_forecast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists acceptable inputs and emphasizes the async requirement to avoid timeouts. It also positions the tool for workforce transformation. However, it does not explicitly state when not to use it or name alternative tools for different use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

logistics_esg_incident_trackerA
Read-onlyIdempotent
Inspect

Tracks real-time ESG incidents in logistics networks for COOs, including supply chain disruptions, regulatory violations, and sustainability risks. Inputs: geographic region, incident type (e.g., emissions, labor, deforestation), and time range. Outputs: structured incident data with severity, location, and source verification. Uses CDP open data and UNCTAD STAT for comprehensive coverage. Keywords: ESG, logistics, supply chain, sustainability, compliance, risk management.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
regionYesGeographic region filter (e.g., 'Europe', 'Asia', 'Global')
endDateNoEnd date for incident search (ISO 8601)
severityNoMinimum severity level to include
startDateNoStart date for incident search (ISO 8601)
incidentTypeYesType of ESG incident to track

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
summaryNo
warningsNo
incidentsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, and idempotentHint=true, indicating safe, non-destructive operations. The description adds value by specifying data sources (CDP, UNCTAD STAT) and output structure (severity, location, source verification), enhancing transparency beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with three focused sentences and keywords. It front-loads the purpose and efficiently covers inputs, outputs, and data sources. No redundant or vague language.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (though not shown), the description adequately explains the tool's purpose, inputs, outputs, and data sources. It covers key aspects for a tracking tool, though it omits the async parameter's polling mechanism, which is documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds semantic context by listing inputs (geographic region, incident type, time range) and providing examples of incident types (emissions, labor, deforestation), which complements the schema without contradicting it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool tracks ESG incidents in logistics networks, specifying the resource (ESG incidents) and scope (logistics networks). It differentiates from siblings by focusing on real-time tracking and specific examples like supply chain disruptions and regulatory violations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for COOs tracking logistics ESG incidents, but lacks explicit guidance on when to use this tool over alternatives like esg_audit_multi or supplier_esg_audit. No exclusion criteria or comparative context provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ma_arbitrage_hunterA
Read-onlyIdempotent
Inspect

As a CFO, identify cross-border M&A arbitrage opportunities by comparing target company valuations across different jurisdictions. Inputs include target company ticker, primary and secondary jurisdictions, and valuation metrics. Outputs include valuation gaps, FX-adjusted multiples, and jurisdiction-specific premiums/discounts. Uses real-time ECB FX rates, Yahoo Finance market data, and SEC EDGAR filings for public companies. Ideal for quick assessment of potential arbitrage in M&A scenarios.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
sectorNoIndustry sector for peer comparison (e.g., 'Technology')
targetTickerYesTarget company ticker symbol (e.g., 'AAPL')
valuationMetricNoValuation multiple to use for comparison
primaryJurisdictionYesPrimary jurisdiction for valuation comparison (e.g., 'US')
secondaryJurisdictionNoSecondary jurisdiction for valuation comparison (e.g., 'DE')

Output Schema

ParametersJSON Schema
NameRequiredDescription
fxRateNo
statusYes
sourcesNo
warningsNo
valuationGapNo
peerMultiplesNo
targetCompanyNo
primaryValuationNo
secondaryValuationNo
jurisdictionPremiumNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds value by disclosing data sources (ECB FX rates, Yahoo Finance, EDGAR) and output specifics (valuation gaps, FX-adjusted multiples). No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four well-structured sentences, front-loaded with purpose. Every sentence adds value: audience, inputs, outputs, data sources, and use case. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of cross-border M&A arbitrage, the description covers purpose, inputs, outputs, and data sources. It is adequate for a quick assessment tool. Could mention prerequisites (e.g., company must be public) but not essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description summarizes inputs but adds no new detail beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: identifying cross-border M&A arbitrage opportunities by comparing valuations across jurisdictions. It specifies the verb (identify), resource (arbitrage opportunities), and scope (cross-border). The purpose is distinct from sibling tools, though not explicitly differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for quick assessment of M&A arbitrage, but lacks explicit guidance on when not to use or comparison with alternatives. It says 'Ideal for quick assessment' which gives some context but no exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ma_deal_screenerC
Read-only
Inspect

M&A Deal Screener — Gapup agent-payable C-suite expertise (CSO). Returns a structured, audited deliverable. Reference case: Salesforce M&A targets — 12 cibles screened · fit score + valuation + integration risk. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
acquirerYes
criteriaYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds that inputs are validated server-side and the tool returns an audited deliverable, which is basic behavioral context but does not go beyond what annotations imply. No contradictions found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short (3 sentences including the reference case), but the first sentence is cryptic ('Gapup agent-payable C-suite expertise (CSO)') and may confuse. It could be streamlined and front-loaded with clearer purpose. The reference case is helpful but adds length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should describe the deliverable's fields. It mentions 'fit score + valuation + integration risk' only in the reference case, not as guaranteed output. Input parameters are not explained, and the async behavior is only in the schema. The description is incomplete for a tool with nested inputs and no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only 'async' parameter has a description in the schema). The tool description does not explain the acquirer or criteria parameters, nor their nested fields. The phrase 'send the documented case fields' is vague. The description adds little semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly indicates that the tool screens M&A deals for an acquirer, providing a structured deliverable with fit score, valuation, and integration risk. The reference case (Salesforce M&A targets) helps clarify the purpose. While not as explicit as 'screens potential targets', the verb and resource are clear. It distinguishes from siblings like re_deal_screener by mentioning M&A.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lacks any guidance on when to use this tool versus alternatives. It does not mention when not to use it or provide criteria for choosing among siblings (e.g., re_deal_screener, ma_arbitrage_hunter). The phrase 'Gapup agent-payable C-suite expertise' is vague and does not clarify context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manufacturing_esg_compliance_mapperA
Read-onlyIdempotent
Inspect

As a COO, quickly identify ESG compliance gaps across manufacturing facilities using EPA TRI emissions data and GRI sustainability standards. Input facility identifiers or geographic regions to receive a prioritized remediation roadmap with risk scores, regulatory violations, and suggested corrective actions. Ideal for sustainability reporting, regulatory risk assessment, and operational improvement planning. Keywords: ESG compliance, manufacturing facilities, EPA TRI, GRI standards, sustainability reporting, regulatory risk.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNoReporting year (default: current year - 1)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
regionNoGeographic region (state, county, or ZIP code) for facility search
includeGriNoInclude GRI standards analysis (default: true)
facilityIdsYesList of EPA facility identifiers (e.g., TRIFID)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
summaryNo
warningsYes
facilitiesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the description's burden is reduced. The description adds context about using EPA TRI and GRI data and producing a roadmap, but it does not mention the async parameter or potential latency, which is a significant behavioral aspect given the async option in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences and a keyword list. It front-loads the target user (COO) and action. The keywords add searchability but slightly clutter. Overall, it is well-structured and succinct.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (ESG compliance with multiple data sources) and the presence of an output schema, the description is fairly complete. It covers data sources, inputs, and outputs. However, it could mention that results may be large or require polling via the async parameter for better completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds minimal meaning beyond the schema—it mentions facility identifiers and regions but does not provide additional context for parameters like year or async. The schema descriptions already adequately explain each parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: identifying ESG compliance gaps across manufacturing facilities using EPA TRI and GRI standards. It specifies inputs (facility IDs or regions) and outputs (prioritized remediation roadmap). However, it does not explicitly differentiate from sibling tools like esg_audit_multi or supplier_esg_audit, which may have overlapping functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for COOs in sustainability reporting and regulatory risk assessment, but it does not provide explicit guidance on when to use this tool versus alternatives (e.g., esg_audit_multi for broader audits). No exclusions or when-not-to-use are stated, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manufacturing_waste_heatmapA
Read-onlyIdempotent
Inspect

Generates manufacturing waste heatmaps for COOs using EPA TRI and FAOSTAT data. Input manufacturing site identifiers or geographic regions to analyze waste streams, emissions, and resource inefficiencies. Outputs include waste intensity maps, circular economy opportunity rankings, and cost-saving potential. Ideal for sustainability strategy and operational efficiency improvements. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearYesAnalysis year (2010-2023)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
regionNoGeographic region (country code or sub-national region) for aggregated analysis
site_idsNoList of manufacturing site identifiers (EPA TRI IDs or FAO facility codes)
waste_typesNoSpecific waste types to analyze (e.g., ['metals', 'chemicals', 'energy'])

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
heatmap_dataNo
opportunitiesNo
benchmark_dataNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint. The description adds that passing async:true avoids timeout, and mentions data sources (EPA TRI, FAOSTAT). It does not contradict annotations, and adds useful behavioral context beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, each conveying essential information: purpose and data sources, input types, outputs, and async usage. No unnecessary words, front-loaded with the main action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (5 parameters, output schema exists), the description covers purpose, inputs, outputs, and async guidance. It does not detail return values (output schema handles that) and could mention optionality of waste_types, but overall adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the description adds little beyond what the schema provides. It mentions 'manufacturing site identifiers or geographic regions' which maps to site_ids and region, but the schema already explains these clearly. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates waste heatmaps using EPA TRI and FAOSTAT data for COOs, and lists specific outputs (waste intensity maps, circular economy opportunity rankings, cost-saving potential). This is specific and distinguishes it from siblings like manufacturing_esg_compliance_mapper.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says 'Ideal for sustainability strategy and operational efficiency improvements', which implies usage context but does not explicitly state when not to use or provide alternatives. With many sibling tools, clearer guidance would help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

margin_doctorC
Read-only
Inspect

Marge par deal — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub — 8 deals pipeline · €28k ARR sous-marge détecté · Récupération €4.2k/an · Playbook 4 scénarios. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
dealsYes
companyYes
productYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. Description adds that inputs are validated server-side and that it returns an audited deliverable, but does not disclose error handling, rate limits, or what happens on validation failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short but includes a lengthy reference case that may not be useful for understanding the tool's purpose. It could be more focused on functional description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested objects, 4 parameters, no output schema), the description fails to explain the return format, what the deliverable contains, or how to interpret results. It leaves significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 25%, and the description only says 'send the documented case fields' without explaining the parameters (company, product, deals) beyond their schema definitions. It does not compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns a structured, audited deliverable related to margin per deal, but it lacks a clear verb (e.g., 'analyze', 'calculate') and uses jargon. It distinguishes from 'margin_doctor_finance' only by name, not explicit differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs. alternatives (e.g., 'margin_doctor_finance'). The description includes a reference case but no explicit when-to-use or when-not-to-use instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

margin_doctor_financeC
Read-only
Inspect

Médecin des Marges — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: Alan — ARR €60M · marge brute 68% → 79% · €3,2M fuites identifiées · Rule of 40 : 14→38. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
costBreakdownYes
marginTargetsYes
unitEconomicsYes
incomeStatementYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark it as readOnlyHint=true and openWorldHint=true. The description confirms it returns a deliverable (no mutation), which aligns. However, it does not disclose any further behavioral traits such as authentication needs, rate limits, or what happens with invalid inputs beyond server-side validation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description includes a reference case which may be useful but occupies space. It could be more concise by focusing on core purpose and parameter usage rather than a marketing-like example.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested objects, no output schema), the description is insufficient. It does not explain the return format, the audit process, or how results are structured. The reference case gives some context but lacks completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has low documentation coverage (17%) and the description adds no explanation for parameters beyond 'send the documented case fields'. It does not describe the meaning of company, incomeStatement, costBreakdown, etc., leaving the agent to infer from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool returns a structured, audited deliverable for financial margin analysis, and provides a concrete reference case (Alan). However, it does not distinguish itself from the sibling tool 'margin_doctor', causing potential ambiguity about the specific role of 'finance' in its name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like 'margin_doctor' or other financial analysis tools. The description only instructs to 'send the documented case fields', which is procedural rather than selective.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

market_entry_strategistB
Read-only
Inspect

Stratégie d'entrée marché — Gapup agent-payable C-suite expertise (CSO). Returns a structured, audited deliverable. Reference case: OpenAI Inde 2026 — entrée marché 1.4Md utilisateurs · 5 forces Porter + 4 entry modes + 18-month roadmap + risk register. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
preferencesYes
targetMarketYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, openWorldHint) already indicate the tool is read-only and generates content. The description adds that the output is 'audited' and that inputs are 'validated server-side', providing minor behavioral context beyond annotations. No contradictions are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (3 sentences) and front-loaded with the tool's purpose. Each sentence contributes: name, output type, reference example, and validation note. The reference case is specific but not overly verbose. Minor deduction for extraneous detail (the exact number of users) that could be generalized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested inputs, no output schema), the description is incomplete. It does not describe the return format beyond 'structured, audited deliverable' or the components listed in the reference case. Without an output schema, the agent lacks full understanding of what to expect, especially for async usage and result polling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'async' has a description). The tool description does not explain any parameter meanings, despite offering a reference case that might imply structure. For a tool with 5 complex parameters (including nested objects), this is insufficient guidance for an AI agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a market entry strategy generator, specifying verb 'returns' and resource 'structured, audited deliverable', with a concrete reference case (OpenAI India 2026) that outlines the analytical framework (Porter's 5 forces, entry modes, roadmap, risk register). This distinguishes it from sibling tools by emphasizing C-suite expertise and audit quality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lacks explicit guidance on when to use this tool versus alternatives. It mentions 'Gapup agent-payable C-suite expertise' but does not specify conditions, exclusions, or when not to use it. No comparison to sibling tools like geographic_expansion or market_sizing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

marketing_roi_dashboardC
Read-only
Inspect

Dashboard ROI marketing — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Gapup Hub — H1 2026 · 5 canaux · ROI 3.2× · Attribution W-shaped · Budget €60k. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
arpuEurYes
channelDataYes
companyNameYes
periodLabelYes
totalRevenueAttribEurYes
targetAttributionModelYes
currentAttributionModelYes
totalMarketingBudgetEurYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side, which is useful but does not disclose other behavioral traits. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief but contains a reference case that may not be universally helpful. It is not overly long, but the jargon and lack of structure reduce clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 9 parameters, nested arrays, and no output schema, the description is insufficient. It does not explain the deliverable's structure, how to interpret ROI, or prerequisites for inputs like attribution models.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11%, and the description does not explain any parameters beyond 'send the documented case fields'. No added meaning for the 9 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description mentions 'Dashboard ROI marketing' and references a case with ROI, channels, and attribution, implying it computes marketing ROI. However, it does not explicitly state the core function, using vague phrases like 'Returns a structured, audited deliverable'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. It mentions 'C-suite expertise (CMO)' but does not compare with any sibling tools, leaving the agent without decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

market_research_briefA
Read-only
Inspect

Generate a structured, sourced market research brief on any market, sector or industry. Returns a machine-readable note with six sections: an executive overview, a market-size estimate (with assumptions and sources — no invented figures), key players, demand & technology trends, risk factors, and a traceable source list. When to use this tool: an agent needs to assess a new market, validate a business opportunity, prepare a pitch, or benchmark a sector before a strategic decision. Data is assembled live from keyless public sources: Wikipedia (sector context), World Bank (macro GDP/population for market sizing), REST Countries (geo context). Fields that cannot be sourced are marked 'unavailable' rather than estimated. Inputs: topic (required), geo and sector (optional refinements).

ParametersJSON Schema
NameRequiredDescriptionDefault
geoNoOptional geography to scope the brief (country name, region, or continent — e.g. 'France', 'Southeast Asia')
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
topicYesMarket or sector to research (e.g. 'electric vehicle batteries', 'B2B SaaS CRM Europe', 'telemedicine Africa')
sectorNoOptional parent sector to disambiguate the topic (e.g. 'healthcare', 'energy', 'software')

Output Schema

ParametersJSON Schema
NameRequiredDescription
geoYes
risksYes
topicYes
sectorYes
trendsYes
sourcesYesAll sources consulted, with URL and retrieval status
overviewYesExecutive summary of the market
key_playersYes
generated_atYesISO-8601 timestamp of generation
market_size_estimateYesMarket size estimate with hypotheses. All figures sourced or marked unavailable.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, openWorldHint, destructiveHint, and idempotentHint. The description adds valuable behavioral context: specific data sources (Wikipedia, World Bank, REST Countries), the 'keyless' public source nature, and the policy of marking unsourced fields as 'unavailable'. It does not discuss rate limits or response time, but the async parameter addresses timeouts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four concise sentences, front-loading the main purpose and output structure, then usage guidelines, then sources and data policy. Every sentence earns its place with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 4 parameters (no enums, output schema exists), the description covers purpose, output sections (six specified), when to use, data sources, and handling of unavailable data. It does not describe return values in detail, but the output schema covers that. Overall, it is complete for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already describes all four parameters. The description adds context by stating topic is required and geo and sector are optional refinements, with examples. The async parameter is omitted from the description but is covered by the schema. This adds marginal value beyond the schema, earning a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it generates a structured, sourced market research brief on any market/sector/industry, listing six specific sections. It distinguishes itself from siblings like 'market_sizing' and 'competitive_deep_dive' by emphasizing the comprehensive, sourced brief format and listing specific use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit 'When to use this tool' section listing four scenarios (assess new market, validate opportunity, prepare pitch, benchmark before strategic decision). It does not explicitly provide when-not-to-use or alternatives, which is a minor gap for a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

market_sizingC
Read-only
Inspect

Dimensionnement marché TAM/SAM/SOM — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Gapup Hub — TAM/SAM/SOM IA décisionnelle C-suite Europe · TAM €48Md · SOM €280M Year-3. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
targetYes
horizonNo
productYes
approachNo
competitorCompsNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=true and openWorldHint=true, indicating a safe read operation with external data access. The description adds 'Returns a structured, audited deliverable' and 'Inputs are validated server-side', which is consistent but does not significantly enhance transparency beyond the annotations. No contradictions are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with a reference case, making it concise but slightly cryptic. The terminology 'Dimensionnement marché' and 'Gapup agent-payable' may confuse some agents. It could be structured more clearly with bullet points or separated sections for input requirements and output format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex input schema (nested objects) and no output schema, the description should provide more context on expected inputs and the structure of the deliverable. The reference case gives numerical examples but lacks explanation of how to supply the required fields. Many sibling tools exist, but the description does not help agents decide when to use this tool over others.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (17%) with only the 'async' parameter described. The description says 'send the documented case fields' but does not specify what fields are required or how they map to parameters. For example, it does not explain that 'product' requires name, category, valueProposition, or that 'target' requires geography, segments, customerType. More parameter guidance is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Dimensionnement marché TAM/SAM/SOM' which clearly indicates market sizing for TAM/SAM/SOM. It mentions returning a structured deliverable and provides a reference case, making the purpose clear. However, it does not explicitly differentiate from sibling tools like market_entry_strategist or competitive_deep_dive, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives. It mentions 'Gapup agent-payable C-suite expertise (CMO)' and 'send the documented case fields' but does not provide explicit context or exclusions. This is insufficient for effective tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ma_tax_efficiency_mapperA
Read-onlyIdempotent
Inspect

For CFOs evaluating cross-border M&A deals: analyzes tax efficiency by mapping withholding tax rates, transfer pricing regulations, and permanent establishment risks across specified jurisdictions. Inputs include acquirer/target jurisdictions, deal structure, and transaction value. Outputs jurisdiction-specific tax exposure, efficiency scores, and risk flags. Uses World Bank Tax Rates API, IMF SDR data, and SEC EDGAR filings for corporate tax disclosures. — pass async:true REQUIRED to avoid x402 timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
deal_structureNoType of M&A transaction structure
transaction_valueNoDeal value in USD millions
target_jurisdictionYesISO 3166-1 alpha-3 country code of the target entity
acquirer_jurisdictionYesISO 3166-1 alpha-3 country code of the acquiring entity
include_transfer_pricingNoWhether to analyze transfer pricing risks

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
tax_treatiesNo
efficiency_scoreNo
target_tax_ratesNo
acquirer_tax_ratesNo
transfer_pricing_riskNo
permanent_establishment_riskNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and idempotent behavior. The description adds value by mentioning the use of external APIs (World Bank, IMF, SEC) and the requirement to pass `async:true` to avoid timeout, providing behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a front-loaded purpose, followed by inputs, outputs, data sources, and a key usage note. Every sentence is informative and there is no redundancy or fluff, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, 2 required), the description provides a solid overview of purpose, inputs, outputs, and data sources. The existence of an output schema reduces the need to detail return values. However, it lacks information on error handling or rate limits, which would enhance completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage, so the description adds marginal value. It lists some parameters (acquirer/target jurisdictions, deal structure, transaction value) but does not elaborate beyond the schema. The async parameter receives extra guidance about timeout, but overall, parameter semantics are adequately covered by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: analyzing tax efficiency for cross-border M&A deals by mapping withholding tax rates, transfer pricing regulations, and PE risks. It specifies inputs and outputs, making the purpose distinct from general tax tools, though it does not explicitly differentiate from sibling tools like ma_deal_screener or tax_compliance_multi.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description targets CFOs evaluating cross-border M&A deals, implying usage context. However, it does not provide explicit guidance on when not to use this tool or suggest alternative tools, leaving the agent to infer usage without clear boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

meddic_scoringC
Read-only
Inspect

Scoring MEDDIC du pipeline — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub — Pipeline 8 deals · €2.1M · MEDDIC score moyen 62/100 · 3 deals at-risk. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
dealsYes
companyYes
productYes
salesCycleNo
targetWinRateNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, so the safety profile is clear. The description adds that it returns an 'audited deliverable', but lacks details on rate limits, authentication, or other behavioral traits. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is short (two sentences) and to the point. It includes a reference case that may be helpful. However, the structure could be improved by front-loading the purpose more clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite complex input schema (nested objects, 6 parameters) and no output schema, the description provides minimal context. It does not explain output format, error handling, or async usage beyond what schema already describes. Overall, incomplete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%, meaning most parameters lack descriptions. The tool description does not compensate; it only mentions 'documented case fields' without adding meaning to individual parameters. Baseline should be higher given low coverage, but description fails to help.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly identifies the tool's purpose: scoring MEDDIC pipeline and returning a structured deliverable. The verb 'scoring' and resource 'MEDDIC pipeline' are explicit. However, it does not distinguish from sibling tools like 'deal_coach' or 'sales_pipeline_forecast' that may overlap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. The description mentions server-side validation but does not specify prerequisites, context, or exclusions. Siblings list suggests many related tools, but no differentiation is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

model_behavior_drift_monitorA
Read-onlyIdempotent
Inspect

Monitors AI model output drift by comparing current model responses against MLCommons safety benchmarks. Designed for risk and compliance personas to detect behavioral deviations that may indicate safety or alignment issues. Accepts model outputs or identifiers and returns structured drift metrics with statistical significance. Sources data from MLCommons public benchmark APIs.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
thresholdNoDrift threshold for alerting
currentOutputsNoRecent model outputs to analyze for drift
baselineMetricsNo
modelIdentifierYesUnique identifier for the model being monitored

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
driftMetricsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly, openWorld, and idempotent hints. Description adds context about sourcing data from external MLCommons APIs, returning structured drift metrics with statistical significance, and accepting model outputs or identifiers. This goes beyond annotations but could mention latency implications of external API calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no waste. First sentence states primary function, second adds persona and outcome, third describes inputs/outputs and data source. Key information is front-loaded and each sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given good annotations, high schema coverage, and presence of an output schema, the description covers purpose, inputs, output type, and data source. It does not mention async behavior or error handling, but these are in schema. Overall complete for a monitoring tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already provides detailed descriptions for all parameters (80% coverage), including async, threshold, currentOutputs, baselineMetrics, and modelIdentifier. Description does not add new parameter semantics beyond mentioning 'model outputs or identifiers' which aligns with schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool monitors AI model output drift by comparing against MLCommons safety benchmarks, targeting risk and compliance personas. It distinguishes from siblings like bias_amplification_tracker or hallucination_confidence_meter by specifying the benchmark source and drift detection focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description implies usage for drift detection with MLCommons benchmarks and risk/compliance personas, but does not explicitly state when to use this tool versus alternatives like bias_amplification_tracker or safety_guardrail_breach_analyzer. No when-not or alternative scenarios are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

model_safety_certification_checkerA
Read-onlyIdempotent
Inspect

Verifies AI model safety certifications against MLCommons and IEEE 7000 standards. Designed for risk management personas to assess model compliance with established safety benchmarks. Accepts model identifiers or certification IDs and returns structured verification results with source references.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
model_idYesUnique identifier for the AI model
standardNoSafety standard to check against
certification_idNoSpecific certification ID to verify

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
complianceNo
last_verifiedNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds context beyond annotations by stating it returns structured verification results with source references. Annotations already declare readOnlyHint, idempotentHint, openWorldHint, so no contradiction. Description supplements well.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences with no fluff. Front-loaded with core action, then purpose, then behavior. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, covering return values. Description adds context about source references and persona. Could mention async behavior, but schema covers that. Overall sufficient for complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the description adds minimal parameter-specific detail. It mentions 'model identifiers or certification IDs' corresponding to model_id and certification_id, but does not exceed schema info. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Explicitly states it verifies AI model safety certifications against MLCommons and IEEE 7000 standards, distinguishing it from sibling tools. The verb 'verifies' and resource 'safety certifications' are specific and clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Identifies target persona (risk management) and input types, but does not provide explicit guidance on when to use this tool versus alternatives or when not to use it. No exclusions or comparisons to siblings given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

monte_carlo_portfolioA
Read-only
Inspect

Pure-compute Monte Carlo portfolio simulation using Geometric Brownian Motion (GBM). Models a multi-asset portfolio across time with contributions, withdrawals, and annual rebalancing. Returns full probability distribution of terminal wealth, percentile paths, drawdown stats, and Sharpe ratio. Modes: simulate (full Monte Carlo) | glide_path (lifecycle 110-age target-date allocation) | stress_test (4 historical crises: 2008 GFC / 2000 dotcom / 1970s stagflation / 2020 COVID). No external data needed — all computed from asset assumptions. Ticker defaults built-in: SPY/VOO/VTI 7%/15%, QQQ 9%/20%, TLT/BND 3%/6%, GLD 5%/18%, BTC 30%/70%. ICP: asset managers, family offices, retail wealth advisors, robo-advisor agents, retirement planners. 10k simulations × 30 years runs in <3s on V8 JIT.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYessimulate = full Monte Carlo GBM | glide_path = lifecycle target-date allocation | stress_test = 4 historical crisis scenarios
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
assetsYesPortfolio assets. Weights must sum to 1.0 (auto-normalized if not).
simulationsNoNumber of Monte Carlo simulations (1000-100000). Default 10000.
horizon_yearsYesInvestment horizon in years (1-50).
target_value_eurNoTarget terminal portfolio value in EUR. Used to compute probability_target_achieved.
confidence_intervalsNoPercentiles to compute in the output distribution. Default [5, 25, 50, 75, 95].
initial_investment_eurYesInitial capital in EUR (e.g. 100000 for €100k).
withdrawals_annual_eurNoAnnual withdrawal amount in EUR for decumulation phase (e.g. 50000 for €50k/yr).
contributions_annual_eurNoAnnual contribution in EUR (e.g. 12000 for €1000/month).
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and openWorldHint, which the description reinforces by stating 'no external data needed.' It adds behavioral details like the computational speed (<3s for 10k simulations × 30 years) and the use of default ticker assumptions, going beyond annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph but packs significant detail without redundancy. It could benefit from more structured formatting (e.g., bullet points for modes), but it remains concise and front-loads the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 parameters, 3 modes, no output schema), the description covers the core functionality, performance, target users, and default values. It does not detail the exact output structure but mentions key outputs (distribution, percentile paths, drawdown stats, Sharpe ratio). More detail on error conditions or exact output format would improve completeness, but it is largely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds value by explaining default ticker assumptions (SPY 7%/15%, etc.) and the meaning of each mode (simulate, glide_path, stress_test with specific crisis scenarios), which are not fully detailed in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description precisely states it is a Monte Carlo portfolio simulation using GBM, with clear modes and output types. It distinguishes itself from sibling financial tools by specifying its unique functionality and target users.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description identifies the ideal customer profile (asset managers, family offices, etc.) and enumerates three distinct modes (simulate, glide_path, stress_test), providing context for when to use each. However, it does not explicitly state when not to use this tool or contrast with alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mttr_breakdown_analyzerA
Read-onlyIdempotent
Inspect

As a CTO, analyze your team's incident response efficiency by breaking down Mean Time To Recovery (MTTR) into root causes: code defects, infrastructure failures, or process bottlenecks. This tool ingests GitHub issue and pull request data alongside Snyk vulnerability reports to provide a detailed breakdown of MTTR components, helping you identify systemic weaknesses in your incident resolution pipeline. Input your GitHub repository details and time range to receive a structured analysis of MTTR contributors with actionable insights.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoYesFull GitHub repository name (owner/repo)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
sinceYesStart date for analysis (ISO 8601)
untilYesEnd date for analysis (ISO 8601)
snykTokenNoSnyk API token for vulnerability data (optional)
githubTokenYesGitHub personal access token for API access

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
breakdownNo
topContributorsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description confirms it is a read-only analysis tool that ingests external data (GitHub, Snyk). It adds context about data sources and output format beyond annotations, e.g., 'structured analysis of MTTR contributors with actionable insights'. No behavioral contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences. It is front-loaded with purpose, then outlines data sources and output. No fluff or repetition. Every sentence adds value: target user and action, inputs, and output.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 6 parameters (with 100% schema coverage) and an output schema (not shown), the description provides a good overview of the tool's functionality and data sources. However, it does not mention the 'async' parameter or how to use it, which is critical for handling slow operations. The output schema covers return values, so that is not a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (repo, since, until, githubToken, snykToken, async). The description reinforces the purpose of repo and time range but does not add new meaning beyond the schema. The 'async' parameter is not mentioned in the description, but the schema adequately explains it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'analyze your team's incident response efficiency by breaking down Mean Time To Recovery (MTTR) into root causes'. It specifies the verb (analyze/break down), resource (MTTR by root causes), and target user (CTO). This distinguishes it from sibling tools like 'dora_metrics_deep_dive' which focuses on broader DORA metrics, and 'incident_response_evidence_collector' which collects raw evidence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage scenarios ('As a CTO, analyze...') but does not explicitly state when to use this tool versus alternatives, nor does it provide exclusions or prerequisites beyond the required inputs. There is no mention of when not to use it, e.g., if you need real-time incident data or have only a single incident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nis2_supply_chain_dependency_mapA
Read-onlyIdempotent
Inspect

Generates a visual dependency map of supply chain relationships under the NIS2 Directive, scoring criticality based on regulatory sources like EUR-Lex and CNIL decisions. Designed for legal and compliance teams to identify high-risk third-party dependencies. Inputs include organization identifiers and optional scope filters. Outputs structured dependency data with criticality scores and regulatory references.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
depthNoDependency chain depth to analyze
scopeNoAnalysis scope: full supply chain or critical dependencies only
sectorNoNIS2 sector classification (e.g., 'energy', 'transport')
organizationIdYesUnique identifier for the organization (e.g., VAT number or LEI)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
dependenciesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint, openWorldHint, and idempotentHint. The description adds behavioral context about generating visual maps, scoring criticality, and referencing regulatory sources (EUR-Lex, CNIL). It does not contradict annotations and provides additional transparency beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the main action, followed by target audience and input/output summary. Every sentence adds value with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich annotations and output schema, the description is complete enough. It explains regulatory sources, criticality scoring, target users, and output structure, covering all relevant context for a generative tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with all parameters described in the schema. The description only summarizes inputs as 'organization identifiers and optional scope filters,' which adds no new meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates a visual dependency map of supply chain relationships under NIS2, scoring criticality based on regulatory sources. It explicitly names the target users (legal/compliance teams) and distinguishes itself from siblings by focusing on NIS2 supply chain dependencies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates it is designed for legal and compliance teams to identify high-risk third-party dependencies, providing clear context. However, it does not include explicit when-not-to-use instructions or alternative tool suggestions, though the sibling list includes many other compliance tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

observability_log_pattern_minerA
Read-onlyIdempotent
Inspect

As a CTO, extract anomalous log patterns from public breach reports (e.g., Verizon DBIR) and MITRE ATT&CK techniques to optimize SIEM rules and observability pipelines. Inputs include threat actor groups, MITRE tactics (e.g., 'TA0005'), or log sources (e.g., 'AWS CloudTrail'). Outputs structured patterns with MITRE mappings, prevalence scores, and detection recommendations. Ideal for reducing false positives and improving breach detection coverage. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
tacticYesMITRE ATT&CK tactic ID (e.g., 'TA0005')
techniqueNoMITRE ATT&CK technique ID (e.g., 'T1059')
log_sourceNoLog source type (e.g., 'AWS CloudTrail', 'Windows Event Log')
max_resultsNo
threat_actorNoThreat actor group name (e.g., 'APT29')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
metadataNo
patternsYes
warningsYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint:true, openWorldHint:true, idempotentHint:true, covering safety and idempotency. The description adds valuable behavioral context: outputs structured patterns with MITRE mappings, prevalence scores, detection recommendations, and a warning about potential timeouts (async hint). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (5 sentences), front-loads the main purpose, and each sentence adds unique value: purpose, inputs, outputs, ideal use case, async tip. No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, annotations, and an output schema, the description covers purpose, inputs, outputs, use case, and async behavior. It does not detail error conditions or rate limits, but annotations handle safety and idempotency, and the async advice suggests potential timeouts. Minor gaps but largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83% (max_results missing description). The description adds context beyond schema by explaining how inputs (threat_actor, tactic, technique, log_source) map to the tool's purpose and giving examples (e.g., 'TA0005', 'AWS CloudTrail'), and clarifies async usage. This compensates for the missing max_results description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool extracts anomalous log patterns from public breach reports and MITRE ATT&CK techniques, which is distinct from sibling tools like observability_metric_anomaly_detector (for metrics) and other security tools. The verb 'extract' and resource 'log patterns from breach reports and MITRE techniques' are specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context on when to use (e.g., 'Ideal for reducing false positives and improving breach detection coverage') and async guidance, but does not explicitly mention when not to use or compare to alternative tools. Sibling list exists but no specific alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

observability_metric_anomaly_detectorA
Read-onlyIdempotent
Inspect

As a CTO, quickly identify anomalous cloud metrics (CPU, latency, memory) by comparing your infrastructure against AWS public benchmarks and CVE-linked hardware risks. Input your observed metrics (e.g., CPU utilization, request latency) and receive a risk assessment with potential root causes. Ideal for performance troubleshooting, security hardening, and capacity planning. Keywords: cloud observability, anomaly detection, CVE hardware risks, AWS benchmark comparison.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
regionNo
metricTypeYes
instanceTypeNo
observedValueYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
cveRisksNo
warningsNo
anomalyScoreNo
benchmarkValueNo
deviationPercentNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and idempotentHint=true. The description adds context by mentioning comparison with benchmarks and CVE risks, implying a non-destructive analysis. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences: first states core purpose, second describes input/output, third lists use cases. No wasted words; front-loaded with key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists (context signals), the description adequately covers the tool's functionality without needing to detail return values. Mentions risk assessment and root causes, providing sufficient completeness for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 20% (low). The description adds meaning by giving examples for metricType (CPU, latency, memory) and observedValue (e.g., CPU utilization). However, it does not explain region or instanceType parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool identifies anomalous cloud metrics (CPU, latency, memory) by comparing against AWS benchmarks and CVE risks. The verb 'identify' and specific resource scope distinguish it from sibling observability tools like 'observability_log_pattern_miner'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context for use cases: 'Ideal for performance troubleshooting, security hardening, and capacity planning.' It does not explicitly state when not to use or name alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

onboarding_salariesC
Read-only
Inspect

Onboarding opérationnel des salariés — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: Pennylane (FR fintech SaaS, ~250 FTE) — 5 parcours 30/60/90 jours · Engineering / Sales / CS / Design / People Ops. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
rolesYes
companyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint true (safe, non-mutating) and openWorldHint true (may use external data). The description adds that it returns a structured, audited deliverable and includes a reference case. It does not disclose specific external data sources or potential side effects beyond what annotations convey, but it provides minimal extra context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief (two sentences plus a reference note) with no fluff. However, the inclusion of specific company names and jargon slightly reduces accessibility. It is front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested objects, 4 parameters, no output schema), the description is incomplete. It does not explain the deliverable's structure, the meaning of the 'focus' parameter, or how to handle asynchronicity. The agent would lack sufficient context to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description contains no explanation of parameters (company, roles, async, focus). Schema description coverage is only 25% (only async has a description). The description fails to compensate for the undocumented 75% of parameters, leaving an agent with insufficient guidance on how to populate inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool is for 'Onboarding opérationnel des salariés' (operational employee onboarding) and returns a structured, audited deliverable. It is specific enough to understand the main function, but jargon (e.g., 'Gapup agent-payable C-suite expertise') slightly obscures clarity. No explicit differentiation from siblings, though no sibling with identical purpose exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description only says 'Inputs are validated server-side — send the documented case fields.' It provides no guidance on when to use this tool vs. alternatives (e.g., comp_benchmark_geo_delta), nor does it mention prerequisites or scenarios where it is not appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

operational_dashboardsC
Read-only
Inspect

Dashboards opérationnels — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: Qonto (5 départements · 12 KPIs) — 4 dashboards live en 3 semaines · time-to-décision -55%. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
techStackYes
departmentsYes
kpiRequestsYes
primaryDashboardToolNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint. Description adds that inputs are validated server-side and returns an audited deliverable, implying a read-only computation. However, no explicit statement about side effects or return formatting beyond 'structured deliverable'. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is two sentences plus a case study, but includes French and jargon ('Gapup agent-payable'). Could be more streamlined for an English agent; front-loading is adequate but wastes space on less essential details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 6 parameters, nested objects, and no output schema, the description fails to explain expected input structure, return format, or how to use results. The case study provides context but is insufficient for complete understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (17%). The description only says 'send the documented case fields' without explaining any specific parameter meanings, leaving agents to guess from schema fields (company, departments, etc.) which lack descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a structured, audited deliverable for operational dashboards targeting C-suite (COO), with a specific reference to Qonto. It distinguishes itself from sibling analytical tools by focusing on operational dashboards, though mixing French reduces clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description only implies usage via a case study, lacking context on prerequisites, exclusions, or sibling tool comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

oss_dependency_velocity_trackerA
Read-onlyIdempotent
Inspect

As a CTO, track the update velocity of your project's open-source dependencies to assess their impact on DORA metrics like deployment frequency and lead time. This tool fetches release history and version adoption data from npm registry and libraries.io, providing insights into dependency freshness, update frequency, and potential risks. Input a list of package names and optional version ranges to analyze. Outputs structured dependency velocity metrics and warnings about stale or rapidly changing packages.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
packagesYes
lookbackDaysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
metricsNo
sourcesNo
warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint. The description adds value by specifying data sources (npm registry, libraries.io) and output type (structured metrics and warnings). It does not disclose potential issues like rate limits or error handling, but the added context is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (3 sentences) and front-loaded with purpose. The opening 'As a CTO' adds role context but is slightly extraneous. Overall efficient with no repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (3 parameters, external data fetch, output schema exists), the description covers purpose, input, and output. It misses detailing the lookbackDays parameter, but the output schema handles return values. Nearly complete for a read-only, idempotent tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 33% (only async described). The description explains the packages parameter ('list of package names and optional version ranges') but fails to mention lookbackDays. This partially compensates for low coverage but leaves a gap for the lookbackDays parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: tracking update velocity of open-source dependencies to assess impact on DORA metrics. It specifies the resource (dependency velocity), verb (track), and differentiates from siblings like dependency_vulnerability_scan by focusing on velocity rather than security.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context (CTO, DORA metrics) and mentions input requirements, but does not explicitly state when to use this tool versus alternatives. Usage is implied but not clearly delineated with when-not-to-use or comparisons to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ossf_scorecard_trend_analyzerA
Read-onlyIdempotent
Inspect

As a CTO, analyze OSSF Scorecard trends for your top 10-50 dependencies to identify security regressions or deteriorating project health. Input GitHub repository names (owner/repo), get structured trend data including score deltas, check failures, and risk flags. Uses OSSF Scorecard API and GitHub Archive for historical context. Ideal for proactive dependency management and risk assessment.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
lookbackDaysNoNumber of days to analyze trends for
repositoriesYesList of GitHub repositories in owner/repo format
minScoreThresholdNoMinimum acceptable score to flag as risky

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
resultsNo
sourcesNo
warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and idempotentHint. The description adds value by naming data sources (OSSF Scorecard API, GitHub Archive) and output structure (score deltas, check failures, risk flags). No contradictions, and it enriches the behavioral model beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences that front-load the role and purpose, then specify input, output, and use case. Every sentence earns its place with zero redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, input format, output structure, and use case. With an output schema present, it does not need to detail return values. Adequate for a tool with moderate complexity (4 parameters, trend analysis).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description offers a high-level explanation of output and input role but does not add detailed parameter meaning beyond what the schema already provides (e.g., lookbackDays, minScoreThreshold).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes OSSF Scorecard trends for dependencies, specifying a verb ('analyze'), resource ('OSSF Scorecard trends'), and scope ('top 10-50 dependencies'). It distinguishes itself from sibling tools, none of which perform similar scorecard trend analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: 'As a CTO' and 'proactive dependency management and risk assessment.' It specifies input format (GitHub repository names). However, it lacks explicit guidance on when not to use this tool or alternatives, missing a higher score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

outbound_sequencerC
Read-only
Inspect

Séquences outbound — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub → CFO + CRO B2B SaaS France — Séquence 6 touches multi-canal · Taux réponse +180%. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
icpYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
offerYes
excludedAnglesNo
targetAccountsNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true, which the description does not contradict. The description adds that inputs are validated server-side and the deliverable is audited, but does not detail other behavioral traits like rate limits, auth needs, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but includes a verbose reference case with metrics that adds little value for tool selection. It could be more concise by focusing on the primary action and output.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 parameters, nested objects, no output schema), the description is insufficient. It does not describe the deliverable's format, content, or how output maps to inputs, leaving agents without crucial context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 20% schema description coverage, the description should compensate. It says 'send the documented case fields' but does not explain the 5 parameters (icp, offer, async, excludedAngles, targetAccounts) beyond what the schema already provides. No additional semantics are added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool returns a 'structured, audited deliverable' related to outbound sequences, but the purpose is vague ('Gapup agent-payable C-suite expertise') and does not clearly articulate the core action or distinguish it from siblings like 'sales_enablement_architect' or 'battle_plan'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is given on when to use this tool versus alternatives. The sibling list is large but the description fails to provide context for appropriate selection or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

partnership_synergiesA
Read-onlyIdempotent
Inspect

Identify and rank strategic partnership opportunities for a company. Returns 5-12 high-fit partnership targets, each scored on revenue lift, time-to-impact, integration complexity and regulatory risk, with a rationale and a recommended first-step outreach playbook. When to use this tool: the user wants business-development or alliance ideas, or M&A target screening before deeper due diligence. Inputs: the user's own company and the strategic axis to unlock through partnership (e.g. enter a new market via distribution, add AI infrastructure without rebuilding). Delivered by Antoine, the AI CSO of the Gapup portfolio.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
constraintsNo
selfCompanyYes
strategicAxisYesWhat strategic axis to unlock through partnership (e.g. 'enter US market via distribution', 'leverage AI infra without rebuild')
currentPartnershipsNoExisting alliances to factor in

Output Schema

ParametersJSON Schema
NameRequiredDescription
kpisNo3-5 headline KPI bubbles
sourcesNo
recommendationsNoPrioritised next steps
executiveSummaryYesBoard-ready partnership opportunity overview
partnershipTargetsYes5-12 ranked partnership targets
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the bar is lowered. The description adds behavioral context by detailing the output structure (scored dimensions, rationale, playbook) and mentions the delivering persona. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core purpose, followed by output, usage, inputs, and persona. Every sentence contributes value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 params, nested objects, output schema), the description provides a clear overview of what it returns and when to use. It covers the strategic context and inputs adequately. Could be enhanced by mentioning data sources or company scope, but overall sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description adds meaning for key parameters: it explains 'strategicAxis' with examples and mentions the company input. However, it does not fully compensate for the low coverage by detailing all parameters or their interactions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verb 'Identify and rank' and resource 'strategic partnership opportunities'. It distinguishes itself from sibling tools by explicitly stating when to use it (business-development/alliance ideas, M&A pre-screening) and implies alternative tools for deeper due diligence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage scenarios: 'When to use this tool: the user wants business-development or alliance ideas, or M&A target screening before deeper due diligence.' It lacks explicit 'when not to use' statements but the context with sibling tools like ma_deal_screener provides implicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patent_landscapeA
Read-only
Inspect

Search, analyze and map patent landscapes across major jurisdictions (US, EP, WO, CN, JP, KR). Three modes: (1) search — find patents by keywords, company name or inventor name; (2) landscape — aggregate distributions: top assignees, top inventors, CPC class breakdown, filings by year, citation leaders, white-space innovation opportunities; (3) lookup — retrieve a specific patent by number (e.g. US10000000B2, EP3456789A1, WO2023/123456). Primary source: WIPO PatentScope (WO PCT, keyless). Optional sources: USPTO PatentsView (US, env PATENTSVIEW_API_KEY), EPO OPS (EP/WO, env EPO_OPS_CONSUMER_KEY + EPO_OPS_CONSUMER_SECRET), Lens.org (global, env LENS_API_TOKEN). Use cases: freedom-to-operate (FTO) analysis, R&D gap identification, VC due diligence IP audit, competitor patent portfolio mapping, inventor network analysis. SLA: <=24s p95 (parallel fetches, 8s per source). Cache: 24h TTL (patent data stable). Quality score: 30 pts per retrieved source (max 90), +10 if >=10 patents, +10 bonus for landscape mode with non-empty top_assignees.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNosearch: keyword/inventor/assignee search; landscape: aggregate distributions; lookup: fetch by patent number. Default: "search"
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesKeywords, company/inventor name, or patent number (e.g. "machine learning", "Tesla Inc", "US10000000B2")
date_toNoISO date YYYY-MM-DD — latest filing date
date_fromNoISO date YYYY-MM-DD — earliest filing date
max_resultsNoMax patents to return (5-50). Default: 20
jurisdictionsNoJurisdictions to include. Default: ["US","EP","WO"]

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
queryYes
statusYes
patentsYes
sourcesYes
landscapeNo
quality_scoreYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and destructiveHint=false. The description adds detailed behavioral information: SLA, cache TTL, quality scoring, source dependencies, and async execution mode, far exceeding annotation expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (modes, sources, use cases, performance) and front-loaded with the main purpose. Every sentence provides useful information without unnecessary repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 params, output schema exists), the description covers modes, sources, performance, and async behavior. It could briefly mention error handling or empty results, but overall is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are already described. The description adds valuable context on modes and sources that enhances understanding, but does not duplicate schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool does patent landscape search, analysis, and mapping across major jurisdictions, listing three distinct modes and common use cases. It is specific and immediately understandable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions explicit use cases like FTO analysis and competitor mapping, but does not directly contrast with sibling tools like patent_ownership_audit or when to avoid this tool. The async parameter is described well.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patent_landscape_asyncA
Read-only
Inspect

Async extended variant of patent_landscape. Supports max_results up to 200 (vs 50 in sync mode) and an optional include_citation_graph flag that enriches each patent with its 2-level citation graph (parent patents that cite this one + child patents cited by this one). Returns immediately (<300ms) with a job_id. Poll the result with patent_landscape_result(job_id) after eta_seconds (~180s). Use for deep R&D white-space analysis, freedom-to-operate (FTO) audits, VC due diligence IP mapping, or large-scale competitor portfolio analysis. Async tool — register a webhook via webhooks_manage(register, url, [job.completed]) to receive callbacks instead of polling. Faster + lighter.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNosearch / landscape / lookup. Default: "search"
queryYesKeywords, company/inventor name, or patent number (e.g. "machine learning", "Tesla Inc")
date_toNoISO date YYYY-MM-DD — latest filing date
date_fromNoISO date YYYY-MM-DD — earliest filing date
max_resultsNoMax patents to return (5-200). Default: 20
jurisdictionsNoJurisdictions to include. Default: ["US","EP","WO"]
include_citation_graphNoIf true, enriches each patent with a 2-level citation graph (parents + children). Adds significant processing time — use for deep analysis only. Default: false.

Output Schema

ParametersJSON Schema
NameRequiredDescription
job_idYesUnique job identifier — pass to patent_landscape_result
statusYes
eta_secondsYes
submitted_atYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description discloses async nature, immediate job_id return (<300ms), eta_seconds (~180s) for polling, and option for webhook callbacks. Annotation readOnlyHint=true aligns with read operation. Adds behavioral context beyond annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with key purpose, then lists features, usage, and options in a logical order. No redundant sentences; each sentence adds substantive information. Appropriate length for complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all essential aspects: async nature, result retrieval, eta, webhook registration, use cases, and parameter highlights. Output schema exists to detail return values, so description doesn't need to repeat that. Complete for a complex async tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, baseline 3. Description adds value by explaining max_results limit (200 vs 50 sync) and citation graph enrichment (2-level). Other parameters like mode and date ranges are adequately described in schema, so minimal additional meaning but still enhances understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it's an async variant of patent_landscape, specifies max_results up to 200 and optional citation graph. Distinguishes from sibling tools 'patent_landscape' (sync) and 'patent_landscape_result' (polling) by mentioning async behavior and alternative use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly lists use cases: deep R&D white-space analysis, FTO audits, VC due diligence, large-scale competitor portfolio analysis. Contrasts with sync variant and provides guidance on polling vs webhook for result retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patent_landscape_resultA
Read-onlyIdempotent
Inspect

Poll the result of a patent_landscape_async job. Returns status=pending while running, status=completed with the full patent landscape report once done, status=failed on error, or status=not_found if the job_id is unknown or expired (TTL 24h). Call this after the eta_seconds hint returned by patent_landscape_async (~180s).

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by patent_landscape_async (prefix: patl_)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds valuable details about status values and the TTL of 24h, going beyond the structured annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all essential: purpose, status behavior, and timing guidance. No filler or redundant information. Front-loaded with the key action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and comprehensive annotations, the description covers the essential behavior. It could mention idempotency (already annotated) or that it's safe to call multiple times, but overall it is complete enough for correct usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes job_id with its prefix (patl_). The description does not add additional semantics beyond implicit reference to the async call. With 100% schema coverage, baseline is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool polls the result of an async job, lists possible statuses (pending, completed, failed, not_found), and explicitly references the async counterpart patent_landscape_async. This distinguishes it from sibling tools like the synchronous patent_landscape and the async submission tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to call after the eta_seconds hint (~180s) from patent_landscape_async, giving clear timing guidance. However, it does not explicitly mention when not to use this tool (e.g., for first-time submission) or compare to alternatives like the synchronous patent_landscape, but context implies it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patent_ownership_auditA
Read-onlyIdempotent
Inspect

Audits patent ownership for employees or contractors, identifying gaps where inventors may not have properly assigned patent rights to the company. Designed for CHROs to ensure IP compliance and mitigate legal risks. Inputs: employee/contractor names or IDs, optional date range. Outputs: list of patents, ownership status, flagged gaps, and assignment details. Sources: USPTO PatFT and EPO Espacenet public records. Keywords: patent audit, IP compliance, employee inventions, contractor agreements, CHRO.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
dateRangeNoOptional date range for patent filings
employeeIdsNoList of employee or contractor IDs (optional if names provided)
employeeNamesYesList of employee or contractor full names to audit

Output Schema

ParametersJSON Schema
NameRequiredDescription
gapsNo
statusYes
patentsNo
sourcesNo
warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, open world, idempotent. Description adds output details (list of patents, ownership status, gaps) and data sources (USPTO/EPO), enhancing transparency without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise, front-loaded with core purpose, then audience, inputs/outputs, sources, and keywords. Every sentence adds value; no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With output schema present, description covers purpose, inputs, outputs, and sources adequately. No apparent gaps for the complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline 3. Description summarizes inputs (names/IDs, date range) but doesn't add new semantics beyond what's in schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool audits patent ownership for employees/contractors and identifies assignment gaps. Distinguishes from siblings like patent_landscape by focusing on ownership and compliance, not general landscape.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly designed for CHROs for IP compliance, with specific sources and keywords. While it doesn't list alternatives, the description provides clear context for when to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

payment_rails_cost_analyzerA
Read-onlyIdempotent
Inspect

As a CFO, compare cross-border payment rail costs (SWIFT, SEPA, local ACH, stablecoins) with FX conversion fees and settlement times. Input source/destination countries and amount, receive cost breakdown, FX rates, and settlement time estimates. Uses ECB FX rates and World Bank remittance price data for accurate cost analysis. Ideal for optimizing international payment strategies and reducing transaction expenses.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
amountYesTransaction amount in source currency
source_countryYesISO 3166-1 alpha-2 country code of payment origin
source_currencyNoISO 4217 currency code of source amount
destination_countryYesISO 3166-1 alpha-2 country code of payment destination
destination_currencyNoISO 4217 currency code of destination amount

Output Schema

ParametersJSON Schema
NameRequiredDescription
amountNo
statusYes
fx_rateNo
sourcesNo
warningsNo
total_costNo
source_countryNo
settlement_timeNo
source_currencyNo
destination_countryNo
destination_currencyNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. Description adds value by specifying data sources (ECB FX rates, World Bank) and outputs (cost breakdown, FX rates, settlement times), which are beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with user role and action, then input/output, data sources, and use case. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Description covers purpose, inputs, outputs, and data sources. It omits potential limitations or error conditions, but for a read-only analysis tool with an output schema, this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear param descriptions. Description mentions source/destination countries and amount, linking them to purpose, but adds no additional semantic depth beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool compares cross-border payment rail costs including FX fees and settlement times, specifying rails like SWIFT, SEPA, local ACH, and stablecoins. It distinguishes from siblings like fx_rate and treasury_optimizer by focusing on cost comparison across multiple rails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description implies usage context: for CFOs optimizing international payment strategies, with input countries and amount. It lacks explicit exclusions or alternatives but provides clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pentest_scope_estimatorB
Read-only
Inspect

Estimateur de scope pentest — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Answers: For a pentest on with assets, what is the effort and cost estimate? · How much should I budget for a web application + API penetration test for SOC 2 Type II compliance? · What is the standard engagement plan (PTES phases + deliverables) for a pentest? · Which engagement type (black-box/grey-box/white-box/red-team) is recommended for my context? · What are the prerequisites and risks for a pentest engagement on my cloud infrastructure? Reference case: Acme SaaS Inc — Fintech B2B EU · web-app + API REST · 12 microservices Node.js AWS · . Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
scope_typeYes
tech_stackYes
asset_countNo
target_geosNo
engagement_typeNo
retest_includedNo
business_contextYes
compliance_frameworksNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that it returns a structured, audited deliverable, but does not disclose additional behaviors such as authentication requirements, rate limits, or what happens on invalid inputs. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the tool's core purpose, but it contains several example questions and a reference case, making it moderately verbose. It could be shortened while retaining key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 parameters, no output schema, many siblings), the description provides a good overview of the tool's function and example uses. However, it lacks details on the return format and expected behavior, which leaves some gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11%. The description adds value by referencing key parameters (scope_type, tech_stack, asset_count, etc.) in example questions, helping users understand the tool's inputs. However, it does not describe each parameter systematically or cover all nine parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it estimates pentest scope, providing effort/cost estimates, engagement plans, and recommendations. It is specific to a niche tool, but does not explicitly differentiate from sibling security tools like cyber_risk_auditor or attack_surface_monitor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Example questions and a reference case imply when to use it, but there is no explicit guidance on when not to use it or mention of alternative tools. The description says 'Answers: ...' which sets clear expectations, but lacks exclusions or context for alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pitch_deck_storylineA
Read-onlyIdempotent
Inspect

Build a complete investor pitch-deck storyline for a company. Returns an 8-20 slide narrative tailored to the target audience (seed-vc / series-a-vc / growth-vc / strategic / bank / grant) — each slide carrying a title, key points, a speaker note and a visual hint — plus a Q&A bank of 10-15 likely board questions and traps to avoid. Output is deck JSON ready to export to Google Slides, Notion or Pitch.com. When to use this tool: the user is preparing a fundraise, a board meeting, or an investor presentation. Inputs: the company profile and the target audience type. Delivered by Sarah, the AI Fundraising lead of the Gapup portfolio.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
audienceYesTarget audience — adapts tone + emphasis + Q&A bank
keyFactsYesHard facts to weave into the deck (traction numbers, milestones, awards)
slideCountYes12 = standard VC deck, 15 = bank-friendly with annexes, 20 = growth/strategic

Output Schema

ParametersJSON Schema
NameRequiredDescription
kpisNo3-5 headline KPI bubbles surfaced from keyFacts
slidesYes8-20 slide objects ready to export to Google Slides / Notion / Pitch.com
qaBanksYes10-15 anticipated investor questions with recommended answers
recommendationsNoFundraising preparation actions
executiveSummaryYesOne-paragraph elevator pitch distilled from the deck
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds behavioral context: generates a narrative, returns slide structure with key points, speaker notes, visual hints, and a Q&A bank, and outputs JSON ready for export. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is 4-5 sentences, front-loaded with purpose, then output details, usage, and inputs. Every sentence adds value; even the trailing 'Delivered by Sarah...' is a brief signature. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 params, nested object, output schema exists), the description covers the main use case, inputs, output format (slide structure, Q&A bank), and when to use. It omits the async parameter but the schema covers it. Overall, it provides sufficient context for an agent to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80% (4 of 5 parameters described in the description: company, audience, slideCount, keyFacts; async not described). The description adds semantic value by explaining audience enum values, slideCount ranges (e.g., '12 = standard VC deck'), and that keyFacts are hard facts. This goes beyond the schema's own descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Build a complete investor pitch-deck storyline for a company.' The verb 'build' and resource 'pitch-deck storyline' are specific. Among sibling tools like 'funding_hunter' and 'investor_shortlist', this tool is uniquely focused on creating a deck narrative, not finding funding or investors.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'When to use this tool: the user is preparing a fundraise, a board meeting, or an investor presentation.' This provides clear context for use. It does not mention when not to use it or compare to specific alternatives, but the scenario is well-defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

polymarket_eventsAInspect

List live Polymarket events, ranked by volume. An event groups several related markets — use it to discover a topic, then polymarket_markets to price it. Returns title, description, start and end dates, and URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum events (default 20)
includeClosedNoInclude finished events (default false)
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It discloses the ranking (by volume), scoping (live events), and return fields (title, description, start/end dates, URL). It does not explicitly say 'read-only,' but the verb 'List' conveys a safe read operation, and the return field list adds useful context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the main purpose. The first sentence states what it does; the second explains the relationship to sibling tools and lists return fields. Every sentence earns its place with zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey return values; it does, listing title, description, dates, and URL. It also mentions live events and ranking. Minor gaps like the direction of 'ranked by volume' or default limit are covered by the schema, so the description is nearly complete for a simple list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all parameters (async, limit, includeClosed). The description does not need to add parameter details; it is consistent with the schema but adds no extra semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb and resource: 'List live Polymarket events, ranked by volume.' It clearly differentiates from siblings by stating it lists events (groups of markets), not individual markets, and points to polymarket_markets for pricing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage guidance: 'use it to discover a topic, then polymarket_markets to price it.' This gives a clear workflow and names the alternative sibling tool, making it unambiguous when to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

polymarket_marketsAInspect

Query live Polymarket prediction markets, ranked by volume. Returns question, implied probability (0-1, derived from the outcome price), volume, liquidity, end date and URL. Optional free-text filter on the question.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
limitNoMaximum markets (default 20)
queryNoFree-text filter on the market question
includeRawNoInclude Polymarket's original fields (default false)
includeClosedNoInclude settled markets (default false)
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that data is live, ranked by volume, and includes a derived implied probability from outcome price, plus other return fields. However, it does not mention rate limits, data freshness, or async job behavior beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no redundancy. The first sentence states the action and ranking; the second lists outputs and the optional filter. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only query tool with well-documented parameters, the description covers the core purpose, return values, and filter capability. The lack of an output schema makes the return-field list valuable, though operational details like rate limits or pagination are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 5 parameters have complete schema descriptions, so the description adds little parameter-specific value. It reiterates the free-text filter, but the schema already documents that. The description does not clarify parameter interplay beyond schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it queries live Polymarket prediction markets ranked by volume, and lists the specific output fields. This distinguishes it from sibling tools like polymarket_events and kalshi_markets by naming the platform and market type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (for live Polymarket markets ranked by volume) but does not explicitly mention alternatives or when-not-to-use conditions. Sibling tools exist, but no exclusions or comparisons are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

positioning_strategistC
Read-only
Inspect

Stratège de positionnement — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Gapup Hub vs Tableau/Pigment/Looker — Angle de différenciation + 5 piliers messaging + battle plan. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
marketYes
companyYes
productYes
aspirationsNo
competitorsYes
customerPainsYes
currentWeaknessesNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint=true and openWorldHint=true, which are consistent with the description stating it returns a deliverable. The description adds that it is 'audited' but does not elaborate on latency, data usage, or other behavioral traits. With annotations present, the bar is lower; the description adds modest value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat concise but includes a lengthy reference case and French phrases that may confuse non-native speakers. It is front-loaded but not optimally structured for quick scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides a high-level view of the output (differentiation angle, 5 pillars, battle plan) but lacks details on output structure, async handling, and how it differs from similar sibling tools. With no output schema, more context would be beneficial.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low at 13%, yet the description does not elaborate on any parameter meaning. It only instructs to 'send the documented case fields,' failing to compensate for the schema's lack of descriptions. The async parameter is ignored.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a structured, audited deliverable for positioning strategy, including differentiation angle, 5 pillars messaging, and battle plan. It references a specific case (Gapup Hub vs Tableau/Pigment/Looker). However, it does not explicitly distinguish from siblings like brand_builder or competitive_deep_dive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description only provides an input validation note ('Inputs are validated server-side') and mentions a reference case, but not usage context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_influencerB
Read-only
Inspect

Presse & influenceurs — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Agicap (levée Série C €70M) — CP + 12 contacts presse Tier-1 · plan de diffusion 14 jours. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
budgetNo
companyYes
targetMediaYes
announcementYes
targetAudienceYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint and openWorldHint, which are consistent with the description's mention of a structured deliverable. The description adds context about server-side validation and the nature of the output, but does not explicitly address the implications of openWorldHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise and front-loaded with the tool's domain and purpose. The reference case is informative but adds some density. Overall, it is structured adequately for quick scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters, nested objects, no output schema), the description is insufficient. It fails to explain the deliverable's format, the role of each parameter, or the expected output structure, leaving significant gaps for an agent to decide when and how to invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (17%), and the description does not explain any of the parameters (e.g., company, announcement, targetAudience, targetMedia). It only directs to 'documented case fields' without elaboration, leaving the agent to infer parameter meanings from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is about press and influencers, returning a structured deliverable, and gives a reference case. However, it lacks an explicit verb (e.g., 'generate' or 'create'), which slightly reduces specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions 'C-suite expertise (CMO)' and a reference case, implying suitability for high-level PR tasks, but does not specify when to use this tool versus alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pricing_in_dealC
Read-only
Inspect

Pricing en Deal — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Agicap × Groupe Rocher — Deal €38k · stade négociation · contre-offre -30% · 3 scénarios pricing · ROI 12×. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
dealYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
redLinesYes
negotiationContextYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already set readOnlyHint=true and openWorldHint=true, indicating safe read-only behavior. The description adds that inputs are validated server-side, which is minor extra context. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (2 sentences plus a reference case) but includes a foreign language case that adds clutter. It is not front-loaded with a clear purpose statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with complex nested inputs and no output schema, the description is insufficient. It does not explain what the deliverable contains, how to interpret results, or when to call this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, and the description provides no parameter-level detail beyond 'send the documented case fields'. With 5 complex nested parameters, the description fails to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description mentions it returns a 'structured, audited deliverable' related to pricing in a deal, but the purpose is vague. The reference case and jargon ('Gapup agent-payable C-suite expertise (CRO)') obscure rather than clarify the tool's function. It's not immediately clear what verb+resource this tool operates on.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like 'pricing_strategist' or 'deal_coach'. The description does not state when-not to use it or provide decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pricing_strategistC
Read-only
Inspect

Stratège de pricing — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: Vercel Pricing 2026 — 4 tiers + usage metering · 3 scenarios pricing chiffrés · ARPU +28% target. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
competitorsYes
currentPricingYes
valuePropositionYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the description's mention of returning a 'structured, audited deliverable' is consistent with read-only behavior. The description adds context about server-side validation and a reference case, but does not disclose additional behavioral traits like idempotency or latency that could affect invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, fitting in a single paragraph with key information: purpose, target user, output type, and an illustrative example. It avoids redundancy and front-loads the main action. However, it could be split into more digestible sentences or structured with bullets for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested objects, no output schema), the description is insufficient. It does not describe the deliverable's format, content, or how the tool derives pricing recommendations. The reference case hints at outcomes but does not generalize. More complete context would include typical outputs and parameter dependencies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (17%); only the 'async' and 'company.name' fields have descriptions. The description vaguely advises to 'send the documented case fields' but does not clarify the semantics of required vs optional parameters, or the constraints on nested objects like 'competitors' or 'currentPricing'. This leaves the agent with minimal guidance beyond parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a pricing strategist for C-suite (CMO) that returns a structured, audited deliverable. The reference to Vercel Pricing 2026 with specific tiers and ARPU target provides concrete examples. However, it does not explicitly differentiate from the sibling tool 'pricing_in_deal', which may handle pricing at a different granularity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions 'Gapup agent-payable C-suite expertise (CMO)' implying CMO-level usage, but provides no explicit guidance on when to use this tool versus alternatives, nor any exclusion criteria. The statement 'Inputs are validated server-side' is generic and does not help in deciding usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

privacy_compliance_auditC
Read-only
Inspect

Audit conformité vie privée — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: Lemlist SAS — SaaS outreach B2B, transferts UE→US Schrems II, RGPD + CCPA + LGPD + UK GDPR. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
presenterScriptNo
targetFrameworksYes
processingActivitiesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds that inputs are validated server-side and it returns a deliverable, which aligns with readOnlyHint (true) and openWorldHint (true). No contradiction. However, it does not disclose additional traits like rate limits, result format details, or scope of the audit. Annotations already cover safety; description adds moderate context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Relatively short (3 sentences), but contains marketing fluff ('Gapup agent-payable C-suite expertise') that reduces clarity. The reference case is helpful but could be omitted for conciseness. Front-loads the purpose adequately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, and the description does not specify what the structured deliverable contains or how the parameters relate to the audit. For a tool with 6 parameters (3 required) and deeply nested objects, more context is needed. Siblings like 'ai_governance_full_report_async' have similar patterns but this lacks depth.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (17%). The description does not explain any parameter's meaning or usage; it only says 'send the documented case fields.' The input schema has complex nested objects (company, processingActivities, presenterScript) with no hints in the description. This adds no value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The tool clearly performs a privacy compliance audit and returns a structured deliverable, as indicated by 'Audit conformité vie privée' and 'Returns a structured, audited deliverable.' However, it does not explicitly distinguish itself from siblings like 'esg_audit_multi' or 'ai_governance_full_report_async', and the marketing language ('Gapup agent-payable C-suite expertise') adds noise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description implies it is for complex privacy audits (e.g., reference case involving RGPD, CCPA, LGPD), but does not state exclusions or prerequisites. 'Inputs are validated server-side' is a technical note, not usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_mappingB
Read-only
Inspect

Mapping des process opérationnels — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: Decathlon France — process Retour produit en magasin · 1700 magasins · 200 retours/j/magasin · -30 à -50% temps cible. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
processesYes
presenterScriptNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side and it returns a structured, audited deliverable, providing behavioral context beyond the annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences. It front-loads the purpose in the first sentence. While slightly rambling in the second sentence, it is still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description should explain the return value. It mentions a 'structured, audited deliverable' but provides no detail on its contents. It also lacks prerequisites or typical usage context, leaving gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only async parameter has a description). The description does not explain the parameters or their meaning beyond 'send the documented case fields', which is insufficient for the complex nested schema. The description fails to compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool maps operational processes and returns a structured deliverable, which is clear. However, it does not differentiate from sibling tool process_mining, which likely has a similar purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for COO-level operational process mapping but does not explicitly state when to use or when to prefer alternatives. No guidance on exclusions or when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_miningC
Read-only
Inspect

Mining des process — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: Gapup Hub — 4 process · €320k gaspillage identifié · 3 quick wins · 5 automations. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
objectivesYes
companyNameYes
mainSystemsYes
topProcessesYes
employeeCountYes
revenueLostEstimateEurNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds that inputs are validated server-side, which is a behavioral detail, but does not disclose other traits like rate limits or authentication needs. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (two sentences plus reference), but the first sentence is ambiguous (mixed languages and jargon). It front-loads unclear information rather than a clear purpose statement. Could be more concise and structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters, no output schema, and many siblings, the description is notably incomplete. It does not detail the input parameters, output structure, or differentiate from similar tools. The reference case provides a concrete example but not comprehensive context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% with only the 'async' parameter described. The description does not explain any parameters; it only vaguely refers to 'documented case fields'. For a tool with 7 parameters, this is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses 'Mining des process' which is the tool name, but adds 'Returns a structured, audited deliverable' and a reference case, indicating it produces a process mining audit. It clearly states the resource and output type, but does not distinguish from the sibling 'process_mapping' tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'process_mapping'. It only mentions server-side validation, which is a technical detail, not a usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

procurement_okr_esg_alignerA
Read-onlyIdempotent
Inspect

Aligns procurement OKRs with ESG targets for COOs using GRI standards and EU TED procurement benchmarks. Inputs include procurement objectives and ESG focus areas (e.g., carbon reduction, supplier diversity). Outputs structured alignment scores, gap analysis, and actionable recommendations. Essential for COOs integrating sustainability into procurement strategy. Keywords: procurement, ESG, GRI, EU TED, OKR alignment, sustainability metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
esgFocusAreasYes
industrySectorNo
procurementObjectivesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
alignmentScoresNo
recommendationsNo
benchmarkComparisonNo
overallAlignmentScoreNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, open-world. Description adds value by mentioning GRI standards and EU TED benchmarks used, and outputs (scores, gap analysis, recommendations). No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each adding distinct information: purpose, inputs/outputs, target user, keywords. No redundancy; front-loaded with the core function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given output schema exists, description sufficiently covers inputs, outputs, standards, and user. Complexity is moderate and adequately addressed for a niche tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Description covers the two required parameters with examples (e.g., carbon reduction for ESG focus areas). Optional parameters (industrySector, async) are not mentioned. Schema coverage is 25%, so description partially compensates but could be more thorough.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb 'aligns' and resource 'procurement OKRs with ESG targets'. Distinguishes from siblings by specifying COOs, GRI standards, and EU TED benchmarks. No ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states target user (COOs) and context (integrating sustainability into procurement strategy). Does not mention when not to use or list alternatives, but context is clear for a specialist tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

procurement_six_sigma_waste_hunterA
Read-onlyIdempotent
Inspect

Analyzes procurement waste for COOs using Six Sigma DMAIC framework and EU TED tender data. Identifies non-value-added activities, overprocessing, and inefficiencies in procurement workflows. Inputs include procurement category, time period, and organizational unit. Outputs waste classification, cost impact estimates, and process improvement recommendations. — pass async:true REQUIRED to avoid x402 timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
time_periodYesTime period for analysis (e.g., '2023-01-01/2023-12-31')
six_sigma_toolNoDMAIC
include_ted_dataNo
organizational_unitNoSpecific business unit or department (e.g., 'EMEA', 'Global Operations')
procurement_categoryYesSpecific procurement category to analyze (e.g., 'IT hardware', 'facilities')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
ted_data_coverageNo
cost_impact_estimateNo
waste_classificationNo
process_improvement_recommendationsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds value by specifying the use of EU TED tender data, the timeout risk requiring async, and the outputs (waste classification, cost impact, recommendations). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences plus a concise async note. All information is front-loaded with no wasted words. Every sentence contributes to understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema (not shown), and the description covers high-level outputs. With schema coverage and annotations, the description is adequate for a moderately complex analysis tool. Could mention output format but not necessary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (4 of 6 parameters have descriptions). The description mentions key inputs but does not add significant meaning beyond the schema for all parameters. The async warning is an important addition but limited to one parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'analyzes', the resource 'procurement waste', the target audience 'COOs', and the framework 'Six Sigma DMAIC'. It distinguishes itself from sibling procurement tools like procurement_spend_optim by focusing on waste identification and improvement recommendations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context on inputs (procurement category, time period, organizational unit) and includes a critical note about using async to avoid timeout. However, it does not explicitly state when not to use this tool versus alternatives like supplier_esg_audit or procurement_spend_optim.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

procurement_spend_optimC
Read-only
Inspect

Optimisation des achats / Spend strategy — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: Tech SaaS €60M ARR — 200 fournisseurs analysés · 20 leviers chiffrés · -€2.4M opex/an target. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
topSuppliersYes
spendCategoriesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the description's mention of returning a deliverable and server-side validation adds minor context. No contradiction, but no disclosure of outputs beyond 'structured, audited deliverable' or behavior on invalid inputs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences including a reference case, which adds context but is slightly verbose. Mixed French/English may reduce clarity. Could be more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, and the description does not explain what the tool returns beyond 'structured, audited deliverable'. Given the complexity of input parameters (nested objects), the description should include return value details. Incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (20% per context signals). The description only says 'send the documented case fields', failing to explain the meaning of parameters or how to construct the nested objects. The schema itself provides descriptions only for some properties, but the description should compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states 'Optimisation des achats / Spend strategy' and 'Returns a structured, audited deliverable', clearly indicating procurement spend optimization. However, the jargon 'Gapup agent-payable C-suite expertise (CFO)' is unclear and doesn't fully distinguish from sibling procurement tools like procurement_okr_esg_aligner.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. The reference case implies context, but there are no 'do not use' conditions or comparisons to sibling tools. The agent receives no decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

programmatic_attribution_calibratorB
Read-onlyIdempotent
Inspect

For ad_revenue_ops persona: calibrates marketing mix models (MMM) by ingesting OpenRTB impression-level data from FreeWheel Marketplace and other programmatic sources. Accepts model parameters, date ranges, and impression IDs as input, returning structured calibration metrics and attribution adjustments. Useful for improving model accuracy with real-time bidding data and validating revenue attribution across programmatic channels.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
endDateYesEnd date for impression data (ISO 8601)
modelIdYesIdentifier of the MMM model to calibrate
startDateYesStart date for impression data (ISO 8601)
impressionIdsNoList of OpenRTB impression IDs to include in calibration
confidenceThresholdNoConfidence threshold for calibration metrics

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
calibrationMetricsNo
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description uses the verb 'calibrates,' which implies modifying model parameters, but the annotations declare readOnlyHint=true. This is a direct contradiction. No further behavioral context (e.g., about idempotency or side effects) is added beyond the annotations, which are contradicted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise at two sentences, but could be more structured. It front-loads the persona and action, then lists inputs and outputs. Every sentence adds value, but there is minor redundancy (e.g., 'programmatic sources' and 'programmatic channels').

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters, including async) and the annotation contradiction, the description is incomplete. It does not address the async behavior parameter, the idempotency guarantee, or clarify that the operation is read-only despite using 'calibrates'. The output schema exists but is not described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description broadly lists parameter types ('model parameters, date ranges, impression IDs') but does not add meaningful detail beyond the schema's individual parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool calibrates marketing mix models by ingesting OpenRTB impression data, identifies the target persona (ad_revenue_ops), and specifies inputs and outputs. This distinguishes it from sibling tools like retail_media_attribution_bridge through its focus on programmatic channels and MMM.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions it's useful for improving model accuracy with real-time bidding data, but does not explicitly state when not to use it or provide comparisons to alternatives among the many attribution-related sibling tools. Usage context is implied but lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

programmatic_brand_safety_auditorA
Read-onlyIdempotent
Inspect

Evaluates programmatic ad inventory for brand safety risks using IAB Tech Lab's standards and GDPR-compliant tracking methods. Designed for ad revenue operations teams to assess inventory quality before bidding. Inputs include domain, page URL, and optional contextual signals. Outputs a structured brand safety score with risk categorization and compliance warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesFull page URL being evaluated
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
domainYesRoot domain of the inventory (e.g., 'example.com')
categoriesNoOptional IAB content categories for contextual analysis
gdprConsentNoGDPR consent string (TCF v2.0)

Output Schema

ParametersJSON Schema
NameRequiredDescription
flagsNo
scoreNoBrand safety score (0-100)
statusYes
sourcesNo
warningsNo
riskLevelNo
gdprCompliantNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly, idempotent, and openWorld hints. The description adds context beyond annotations by specifying the evaluation standards (IAB, GDPR), the output structure (brand safety score, risk categorization, compliance warnings), and the optional contextual signals. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no fluff. First sentence defines purpose and standards, second sentence targets users and timing, third sentence lists inputs and output. Efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters and an output schema exists, the description covers the main inputs and output structure adequately. It mentions standards and compliance. However, it omits the async behavior (relevant for the async parameter) and does not detail the output schema (but output schema exists, so not required). Overall, fairly complete for a moderately complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description mentions domain, page URL (which maps to domain and url), and optional contextual signals (categories), but does not explain the async parameter or gdprConsent beyond stating GDPR-compliance. It adds some value but not fully compensating for missing details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates programmatic ad inventory for brand safety risks using IAB standards and GDPR-compliant methods. It specifies the target users (ad revenue operations teams) and the timing (before bidding), distinguishing it from sibling tools like privacy_compliance_audit or ugc_moderation_classifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage 'before bidding' but does not explicitly state when to use this tool versus alternatives, nor does it provide exclusions or scenarios where this tool should not be used. Given the abundance of sibling tools, more explicit guidance would be beneficial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proposal_generatorC
Read-only
Inspect

Générateur de propositions commerciales — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Spendesk × Gapup Hub — Proposition 7 sections · ROI 3Y €1.8M · Payback 4 mois. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
offerYes
companyYes
prospectYes
dealContextNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and openWorldHint. The description adds that the tool returns a structured, audited deliverable and validates inputs server-side, but does not reveal other behavioral traits such as rate limits, auth requirements, or what happens on failure. With annotations, the added value is moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief, with two sentences and a reference case. It is front-loaded with the tool's core function. However, the specific reference case may be too verbose for an agent and could be shortened without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters with nested objects and no output schema, the description is incomplete. It lacks details on how to structure the input objects, what the output looks like, and how the tool integrates with its environment. The reference case provides some context but is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, yet the description does not compensate by explaining the parameters or their roles. It simply says 'send the documented case fields' without detailing the nested objects or required fields. This leaves the agent with insufficient guidance to correctly populate the inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it generates commercial proposals for C-suite expertise (CRO) and returns a structured deliverable. It cites a specific reference case, making the purpose concrete. However, it does not explicitly differentiate from sibling tools like 'pitch_deck_storyline' or 'battle_plan', which could overlap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions that inputs are validated server-side and to send documented case fields, but provides no explicit guidance on when to use this tool versus alternatives. No exclusions or context for when not to use it are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

qa_pre_flightC
Read-only
Inspect

Préparation Q&A investisseurs — Gapup agent-payable C-suite expertise (FUNDRAISING). Returns a structured, audited deliverable. Reference case: Agicap Série C €70M — 30 Q&A stratégiques · 8 questions pièges · Plan de préparation 21 jours. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
roundYes
companyYes
founderContextYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint. The description adds that the tool returns a structured deliverable and validates inputs server-side, but does not disclose side effects or other behavioral traits beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately concise but includes a detailed reference case that may not be essential. It is front-loaded with purpose but could be streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested objects, no output schema), the description lacks details on the deliverable's structure and how to interpret the result. Incomplete for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (25%), and the description adds little parameter-specific meaning beyond 'send the documented case fields'. Nested objects and their fields are not explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool prepares investor Q&A for fundraising and returns a structured, audited deliverable. It references a concrete case (Agicap). However, it does not explicitly differentiate from sibling tools like 'audit_pre_flight'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description mentions that inputs are validated server-side but does not provide context for when to choose this tool over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

qbr_autoC
Read-only
Inspect

QBR automatique CSM — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub × Alan — QBR Q1 2026 · Health score 82/100 · Upsell €18k détecté · Renewal low risk. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
winsYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
periodYes
companyYes
metricsYes
customerYes
challengesYes
nextQuarterGoalsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint, indicating non-modifying and open schema. The description adds minimal value by stating inputs are validated and a deliverable is returned, but does not disclose performance, auth needs, or other behavioral traits beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (two sentences plus a reference case). While concise, the reference case adds noise without clarifying functionality. The structure is acceptable but could be streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 8 parameters and no output schema, the description is incomplete. It does not explain the return format, async behavior (despite async parameter), or how to interpret the structured deliverable. Missing critical details for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13% (async described). The description does not elaborate on any parameters, such as the nested objects (company, customer, metrics) or their meaning. With low coverage, the description should compensate but fails to add semantic context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns a structured, audited QBR deliverable, which is clear. However, it does not explicitly differentiate from sibling tools like 'renewal_optimizer' or 'enps_auto', missing an opportunity to clarify uniqueness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description lacks context on prerequisites, when to choose it over similar tools, or exclusion criteria. Only hints at server-side validation but no usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

real_estate_intelA
Read-onlyIdempotent
Inspect

Real estate intelligence aggregator with a best-in-class French dataset (DVF — Demandes de Valeurs Foncières — 100% of FR transactions since 2019, public, keyless) plus UK Land Registry Price Paid (all UK transactions 1995+). Four modes: (1) property — full transaction history for a specific address; (2) comparables — median/std price/m² within a radius (default 500m); (3) market — annual price series, YoY change, volume, trend by commune; (4) valuation — two-method estimate (comparables median + hedonic regression if n≥30) with confidence scoring (high/medium/low). All sources are free and require no API key. ICP: PropTech agents, REITs, fund managers, family offices, insurance. SLA: ≤25s p95 (sources fetched in parallel, 8s budget each). Cache: 24h TTL (DVF data is stable). Quality score: 30 pts DVF retrieved, 20 pts geocoding, 20 pts UK LR retrieved, 15 pts if comparables count ≥10, 15 pts if method quality achieved. Status: failed/<60/≥60 → failed/partial/final. No env vars required.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesproperty: transactions at an address | comparables: sample around a point | market: commune/neighbourhood market stats | valuation: price estimate for a given surface
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
date_toNoISO date YYYY-MM-DD — latest transaction date
locationYesLocation descriptor. One of: {address, city?, country?} | {lat, lon, radius_m?} | {insee_code} for FR communes.
date_fromNoISO date YYYY-MM-DD — earliest transaction date
max_resultsNoMaximum number of results to return (5–50, default 20)
surface_maxNoMaximum surface in m² (±20% tolerance applied for comparables)
surface_minNoMinimum surface in m² (±20% tolerance applied for comparables)
property_typeNoFilter by property type (default: all)

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
marketNomode=market — commune-level market stats
statusYes
sourcesYes
propertyNomode=property — transactions at the location
valuationNomode=valuation — price estimate
comparablesNomode=comparables — aggregated comp stats
quality_scoreYes
location_resolvedYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds significant behavioral context: data source details (DVF, UK Land Registry), SLA (≤25s p95), cache TTL (24h), quality scoring, status levels, and that no env vars are required. This enriches the agent's understanding beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but well-structured with bullet points and clear sections (modes, data sources, SLA, etc.). It is front-loaded with the core value. While dense, every sentence adds value; minor reduction would improve conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 parameters, nested objects, 4 modes, multiple data sources, quality scoring), the description is very complete. It covers data sources, SLA, caching, target users, and status outcomes. An output schema exists, so return values are not needed in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds meaning by detailing each mode's purpose and linking them to the mode parameter, and by providing context like default radius (500m) and surface tolerances (±20%). This goes slightly beyond what the schema alone provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a real estate intelligence aggregator with specific datasets (French DVF, UK Land Registry) and four distinct modes (property, comparables, market, valuation). This provides a specific verb+resource and distinguishes it from sibling tools, which are largely unrelated or different in scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the four modes and their use cases (e.g., 'property — full transaction history for a specific address') and mentions the ICP (target users). However, it does not explicitly state when NOT to use this tool or provide comparisons to alternative tools among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

realtime_data_streamsA
Read-only
Inspect

High-frequency real-time market data for trading agents, market-making bots and fintech analysts. Returns FX ticks (bid/ask/spread), intraday OHLCV candles, crypto orderbook snapshots (depth 5-50), recent trades with VWAP, and sovereign bond yields. All sources are keyless public REST APIs (Binance, Coinbase, Kraken, OKX, open FX feeds, worldgovernmentbonds.com). Ultra-short cache: 10s for ticks/trades, 60s for orderbook. Use when an agent needs live market data as precise numeric inputs for trading logic, arbitrage detection, or portfolio valuation.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesData stream type: fx_tick (latest FX bid/ask/mid/spread), fx_history_intraday (OHLCV candles), crypto_orderbook (order book snapshot), crypto_trades_recent (last 50 trades + VWAP), bond_yields (sovereign yield %)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
depthNoOrderbook depth (levels each side) for crypto_orderbook mode (default: 20)
periodNoCandle period for fx_history_intraday mode (default: 5m)
symbolYesMarket symbol. FX: EURUSD, GBPUSD, USDJPY. Crypto: BTCUSDT, ETHUSDT, BTC-USD. Bonds: US10Y, US2Y, DE10Y, FR10Y, UK10Y, JP10Y, IT10Y

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
symbolYes
fx_tickNo
sourcesYes
fx_historyNo
bond_yieldsNo
crypto_tradesNo
quality_scoreYes
crypto_orderbookNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds valuable behavioral details: sources are keyless public REST APIs, cache durations (10s for ticks/trades, 60s for orderbook), and the async parameter behavior (returns job_id if async=true). This exceeds what annotations alone provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that is informative and front-loaded with the tool's purpose. It lists data types, sources, and cache durations efficiently. Could be slightly more concise by removing parenthetical lists, but overall it earns its sentences without excessive verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (has output schema: true), the description does not need to explain return values. It covers data types, sources, caching, async behavior, and use cases comprehensively. No significant gaps are present for a real-time data streaming tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter having a clear description. The tool description does not add significant meaning beyond the schema but provides context for modes and symbols. Baseline score of 3 is appropriate since the schema already does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides high-frequency real-time market data for trading agents, market-making bots, and fintech analysts. It enumerates specific data types (FX ticks, OHLCV candles, crypto orderbook snapshots, trades with VWAP, bond yields) and sources, distinguishing it from sibling tools that cover other financial analysis tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'Use when an agent needs live market data as precise numeric inputs for trading logic, arbitrage detection, or portfolio valuation.' It provides context but does not mention when not to use it or list alternative tools, which would strengthen the guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recruiting_architectC
Read-only
Inspect

Architecte du recrutement — Gapup agent-payable C-suite expertise (CHRO). Returns a structured, audited deliverable. Reference case: Stripe France — 12 postes Q3 2026 · sourcing multi-canaux + employer brand + frameworks d'entretien + parcours candidat · time-to-hire -45%. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
rolesYes
budgetYes
companyYes
preferencesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint, which are consistent with the description (no modification, handles varied inputs). The description adds server-side validation and deliverable output, but no further behavioral details like auth or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (two sentences) but includes a case study which adds length without structure. It is concise but could be better organized with key information front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested objects, no output schema), the description is insufficient. It does not explain the deliverable content, how to interpret results, or handle errors, leaving the agent with little guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%, and the description does not explain any parameters beyond 'send the documented case fields.' This fails to add meaning beyond the schema, especially given complex nested objects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a recruitment architecture tool returning a structured deliverable, with a reference case illustrating its use. It distinguishes from sibling tools by focusing on high-level C-suite expertise, but does not explicitly differentiate from other recruiting tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description mentions 'agent-payable C-suite expertise' implying high-level use, but lacks when-not-to-use or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

re_deal_screenerA
Read-only
Inspect

Screener deal immobilier (EU) — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Answers: Screen this real estate deal: , , asking € — give me cap rate vs market, location score, risk flags, and deal recommendation. · Should I pursue this hotel investment at for € with keys? Run an EU deal screener with DVF comparables and Géorisques risk data. · What is the real estate market valuation for a at based on recent French DVF transactions? · Run a due diligence deal screen on this property: , €, sqm — flood risk, cap rate, price vs comparables. · Evaluate this commercial real estate deal for an investment committee: at , €, NOI €. Reference case: Hôtel boutique 45 keys · 12 rue de la Paix 75002 Paris · €12.5M · €277k/key · comp DVF €250-380k/key · location 92/100 · score 72 · pursue-with-conditions. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
addressYes
deal_typeYes
country_iso2YesFR
units_or_keysNo
gross_area_sqmNo
current_noi_eurNo
asking_price_eurYes
investment_thesisNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and openWorldHint=true. The description adds that it returns a structured, audited deliverable and mentions async behavior (returns job_id for slow processing). It also notes server-side validation. These details go beyond the annotations and are consistent with them, providing useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy with multiple example queries and a detailed reference case. While it front-loads the purpose, it becomes wordy. Some information could be condensed without losing clarity. The structure flows from general purpose to examples, but the length impacts conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 parameters, no output schema), the description covers the main use cases, outputs, and data sources (DVF, Géorisques). It mentions the deliverable includes cap rate, location score, risk flags, and recommendation. The async behavior is explained. However, it does not detail optional parameters or parameter interactions, leaving some gaps for complete usage understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11%, meaning most parameters lack schema-level descriptions. The tool description does not systematically explain each parameter's meaning or format. It gives example queries that imply mappings (e.g., 'price' to asking_price_eur, 'keys' to units_or_keys), but these are not explicit. For a tool with 9 parameters, this leaves ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a real estate deal screener for the EU, specifically using French DVF and Géorisques data. It enumerates the output: cap rate vs market, location score, risk flags, and deal recommendation. Example queries demonstrate its scope (hotel, commercial, due diligence). It distinguishes from sibling tools like 'ma_deal_screener' and 'real_estate_intel' by focus on EU real estate with specific data sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides multiple example queries that illustrate when to use the tool: screening deals, evaluating investments, running due diligence, assessing commercial real estate. It references a case study. However, it does not explicitly state when not to use it or contrast with alternatives. The context is clear but lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

renewal_optimizerB
Read-only
Inspect

Optimiseur de renouvellements — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub — Renewals 10 comptes · €89k ARR à 90j · 3 comptes at-risk · Playbook 6 scénarios. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
horizonNo
productYes
accountsYes
targetRenewalRatePctNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, and the description reinforces this by indicating inputs are validated server-side and returning a deliverable. No destructive behavior is implied, so the description adds useful context without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core purpose. However, the reference case example, while illustrative, adds length without essential information and may not be universally applicable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested objects, no output schema), the description is insufficient. It does not explain what the deliverable contains, how to interpret results, or what constitutes a 'documented case'. This leaves significant gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only 'async' has a description in the schema). The tool description does not mention any parameters, so it fails to compensate for the low coverage. The agent receives no additional meaning about inputs beyond the schema's minimal descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a renewal optimizer that returns a structured, audited deliverable. The name and context make the purpose evident, though it does not explicitly differentiate from sibling tools like 'churn_defender' or 'save_plays' that also deal with renewals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description merely instructs to send documented case fields, but does not specify prerequisites, context, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_rate_arbitrage_scannerA
Read-onlyIdempotent
Inspect

Scans for arbitrage opportunities between repo rates (ECB) and short-term funding markets (Treasury Direct). Designed for CFOs to identify cost-effective funding strategies. Inputs include optional date ranges and currency filters. Outputs structured arbitrage opportunities with rate differentials and confidence scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
endDateNo
currencyNo
startDateNo
minDifferentialNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
opportunitiesNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, covering safety and idempotency. The description adds that outputs are structured arbitrage opportunities with rate differentials and confidence scores, which is useful but not critical. There is no contradiction with annotations. The description provides moderate additional context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two sentences: the first clearly states the purpose, and the second adds the target audience and output. There is no unnecessary information, and the key points are front-loaded. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 5 optional parameters, and the description mentions that inputs include optional date ranges and currency filters, but does not cover all parameters (e.g., minDifferential, async). Since an output schema exists, return values are not needed in the description. However, the description is incomplete regarding full parameter semantics and async behavior, which is partially covered by the schema. Overall, it is adequate but has gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'async' has a description). The description mentions 'optional date ranges and currency filters' but does not explain 'minDifferential' or the async behavior detail (though schema covers async). Given the low coverage, the description should compensate by explaining more parameters, but it only glosses over two categories, leaving the other parameters' semantics unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans for arbitrage opportunities between repo rates (ECB) and short-term funding markets (Treasury Direct). The verb 'scans for' and the specific resource (arbitrage opportunities) are precise. Among many sibling tools, this one is distinct due to its focus on repo rates and Treasury Direct, making it easy to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates the tool is designed for CFOs to identify cost-effective funding strategies, implying when to use it. However, it does not provide explicit guidance on when not to use it or compare it to alternatives like 'tariff_arbitrage_finder' or 'ma_arbitrage_hunter', which exist among siblings. The usage context is implied but lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reputation_engineC
Read-only
Inspect

Moteur de réputation — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Reference case: PayShield SaaS — Monitoring réputation Q2 2026. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
brandYes
channelsYes
industryYes
keywordsYes
historicalCrisesNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description aligns by stating it returns a deliverable. It adds context about server-side validation and a reference case. However, it does not disclose additional behavioral traits like response format, pagination, or rate limits beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, but the first part is cryptic ('Gapup agent-payable C-suite expertise (CMO)') and the reference case may be too specific. It is not as concise as it could be, and the structure could front-load the core purpose more clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters, 4 required, and no output schema, the description should explain the deliverable structure and handling of the async parameter. It mentions none of these, leaving critical gaps for an AI agent to correctly invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only async has a description). The tool description does not add meaning to the other 5 parameters (brand, channels, industry, keywords, historicalCrises). It briefly mentions 'send the documented case fields' but does not explain individual parameters, leaving a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool is a reputation engine that returns a structured, audited deliverable, with a specific reference case. The verb 'returns' clarifies the action. However, it does not differentiate from potential sibling tools like sentiment_news_pulse or brand_equity_voice_share_calculator.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives. It mentions 'Gapup agent-payable C-suite expertise (CMO)' which vaguely suggests context but does not specify when-not or provide exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

research_paper_qaA
Read-only
Inspect

Synthèse littérature scientifique (PaperQA2) — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Answers: Conduct a literature review on — what does the evidence show across recent papers? · Evaluate the current hypothesis that — supporting and contradicting evidence with citations. · Map contradictions in the literature on — which camps exist, how many papers per side? · What is the state-of-the-art understanding of as of ? · Perform an interdisciplinary synthesis on — findings from and . Reference case: Gut-brain axis · Cognitive performance in healthy adults · OpenAlex+SemanticScholar+CORE · Evidence synthesis · DOI-verified citations · Contradictions + gaps mapped. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
max_papersYes
year_rangeNo
focus_domainYesall
include_preprintsYes
research_questionYes
evidence_grade_requiredYesstandard
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true, aligning with the description's emphasis on research and evidence synthesis. The description adds context about returning a 'structured, audited deliverable' and server-side validation, but does not disclose additional behaviors like rate limits or authentication requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively long and includes multiple example questions and a reference case, which aids clarity but also adds redundancy. It lacks a concise summary upfront and could be more focused on essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, nested objects, no output schema), the description falls short. It does not explain the return format beyond 'structured, audited deliverable', nor does it describe the behavior of parameters like year_range or evidence_grade_required. The async parameter is only documented in the schema, not in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% (only the 'async' parameter has a description). While the description lists some parameters in examples (research_question, focus_domain, max_papers, etc.), it does not explain their semantics or provide context beyond indicating they are documented fields. This is insufficient compensation for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a literature synthesis tool (PaperQA2) that returns a structured, audited deliverable. It lists specific use cases like literature review, hypothesis evaluation, contradiction mapping, and interdisciplinary synthesis, making its purpose unambiguous and distinct from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit example questions for when to use the tool (e.g., 'Conduct a literature review on <topic>', 'Evaluate the current hypothesis that <claim>'). However, it does not specify when not to use it or mention alternative tools (like 'sci_literature_search'), lacking explicit exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retail_media_attribution_bridgeA
Read-onlyIdempotent
Inspect

Provides unified attribution insights for retail media and programmatic campaigns by analyzing MMM signals from FreeWheel Marketplace and Common Crawl. Designed for ad revenue operations teams to bridge cross-channel performance gaps. Accepts campaign IDs, date ranges, and channel filters as input. Returns structured attribution data with source provenance and confidence scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
endDateYesEnd date for attribution window (YYYY-MM-DD)
channelsNoChannels to include in analysis
startDateYesStart date for attribution window (YYYY-MM-DD)
campaignIdsYesList of campaign identifiers to analyze
confidenceThresholdNoMinimum confidence score for included signals

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
attributionNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, openWorldHint, and idempotentHint, but the description adds valuable behavioral context: it analyzes MMM signals from FreeWheel Marketplace and Common Crawl, and returns structured data with source provenance and confidence scores. This goes beyond the annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences: first states the core function and data sources, second identifies the target audience, third outlines inputs and outputs. Every sentence serves a purpose with no wasted words. Front-loaded with the most critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema and comprehensive annotations, the description sufficiently covers the tool's purpose, input parameters, behavioral context, and output nature. It provides a complete understanding for an AI agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mentions campaign IDs, date ranges, and channel filters but does not add new meaning beyond what the schema already provides for each parameter. No per-parameter elaboration is given in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as providing unified attribution insights for retail media and programmatic campaigns by analyzing specific data sources. It identifies the target audience (ad revenue operations teams) and the core function (bridging cross-channel performance gaps). This is a specific verb+resource combination that distinguishes the tool from generic attribution tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool ('bridge cross-channel performance gaps' for ad revenue operations) but does not explicitly contrast with sibling tools like 'programmatic_attribution_calibrator' or 'retail_media_esg_compliance'. It gives a strong usage scenario without exclusions, meeting the 'clear context' criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retail_media_esg_complianceA
Read-onlyIdempotent
Inspect

Audits retail media networks for ESG compliance by analyzing ad placements, tracking cookies, and verifying ethical advertising standards. Designed for ad_revenue_ops teams to ensure GDPR and sustainability compliance across digital retail platforms. Accepts domain lists or network identifiers as input and returns structured compliance reports with warnings and source references. Requires async:true to avoid timeout errors.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
domainsNoList of retail media network domains to audit
checkESGNoEnable ESG advertising standards compliance check
checkGDPRNoEnable GDPR cookie tracking compliance check
networkIdsNoList of retail media network identifiers

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
summaryNo
warningsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint. The description adds value by disclosing the asynchronous behavior requirement (async:true) and the output format (structured compliance reports with warnings and source references). This goes beyond what annotations provide, without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long with no extraneous information. The first sentence clearly defines the purpose, the second adds context (target users), and the third covers inputs, outputs, and a critical usage note about async. It is well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and annotations, the description covers the main aspects: input types, output format (compliance reports with warnings/references), and the async requirement. However, it could have elaborated on how to use the async mechanism (e.g., polling with job_result) or mentioned any rate limits or data volume concerns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all 5 parameters. The description paraphrases 'domains' and 'networkIds' as inputs but adds no new detail about data types, constraints, or usage beyond what the schema already provides. The statement 'Requires async:true' is slightly misleading as async is optional (no required parameter).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool audits retail media networks for ESG compliance by analyzing ad placements and tracking cookies. It specifies the target users (ad_revenue_ops) and compliance areas (GDPR, sustainability). However, it does not explicitly differentiate from sibling ESG tools like 'esg_audit_multi' or 'action_plan_esg', which could lead to confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context for usage (ensuring GDPR and sustainability compliance for retail media networks) and includes an important note about requiring async:true to avoid timeouts. However, it does not specify when NOT to use this tool or mention alternative tools for similar tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revops_architectC
Read-only
Inspect

Architecte RevOps — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Qonto — ARR €200M · 200 reps · forecast ±35% · fuite €4,2M/an identifiée · plan RevOps 12 semaines. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
keyMetricsYes
objectivesYes
revenueTeamYes
currentStackYes
horizonMonthsYes
currentPainPointsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the description's mention of a 'structured, audited deliverable' aligns with a read-only operation. It adds that inputs are validated server-side, but does not elaborate on other behavioral traits like latency or persistence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences plus a reference case) and front-loaded with the tool's purpose. The reference case adds context convincingly, though the mix of French and English slightly reduces clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested input schema (8 parameters, many nested) and no output schema, the description lacks detail on the deliverable format, return structure, or how to interpret results. The reference case partially compensates but is insufficient for complete agent understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13% (only 'async' has a description). The tool description does not explain any of the required nested fields (e.g., company, keyMetrics), relying solely on the schema. This is insufficient for agent correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool provides a structured, audited RevOps deliverable, with a reference case illustrating its output. However, it does not explicitly define the scope beyond 'C-suite expertise (CRO)', and the purpose could be more specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool over alternatives, such as other architect tools in the sibling list. It only implies usage via the title and reference case, leaving the agent without context for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rfp_tender_architectC
Read-only
Inspect

Architecte d'appels d'offres — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: AO DINUM — Plateforme IA souveraine. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
rfpTypeYes
rfpScopeYes
budgetRangeYes
deadlineISOYes
clientCompanyYes
ourPositioningYes
compliancePointsNo
competitorsLikelyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description's statement about returning a deliverable is consistent. The description adds server-side validation and a reference case, but no additional behavioral traits (e.g., cost, authentication). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short but includes unclear jargon ('Gapup agent-payable C-suite expertise (COO)') that does not earn its place. It could be more concise and focused.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 9 parameters, 7 required, and no output schema, the description is insufficient. It does not explain the deliverable's content, how to interpret results, or provide adequate context for agent understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 11% (rfpType has enum), and the description does not explain any parameter meaning. It simply says 'send the documented case fields' without elaborating, leaving the agent with no help on input semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns a 'structured, audited deliverable' and mentions a reference case, but it does not specify what kind of deliverable (e.g., strategy document, analysis). The jargon 'Gapup agent-payable C-suite expertise (COO)' is unclear. It is not distinguished from sibling tools like proposal_generator or battle_plan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description implies it is for C-suite expertise but does not provide explicit context, exclusions, or comparisons with sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rse_policy_builderC
Read-only
Inspect

Architecte de politique RSE — Gapup agent-payable C-suite expertise (SUSTAINABILITY). Returns a structured, audited deliverable. Reference case: TechCorp SAS — Politique RSE 2025-2028 (500 FTE, €60M CA, SaaS B2B France). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
valuesYes
companyYes
ambitionsYes
targetLabelsNo
currentInitiativesNo
targetStakeholdersYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that it 'returns a structured, audited deliverable' and that inputs are validated server-side, but does not disclose any behavioral traits beyond what annotations imply (e.g., state changes, auth requirements, rate limits). It is consistent with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (two sentences plus a reference) and front-loads the core purpose. No redundant or wasted text. However, it is concise at the cost of completeness for parameters and usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 8 parameters including nested objects and no output schema, the description is insufficient. It does not explain the output format, optional parameters (async, focus, targetLabels, currentInitiatives), or how the deliverable is structured. The reference case helps but does not compensate for missing details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13%, meaning most parameters (including the nested company object, ambitions, etc.) lack descriptions. The description does not add meaning beyond the schema's required fields list. It merely says 'send the documented case fields' without explaining parameter semantics or format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it builds CSR policies (Architecte de politique RSE) and returns a structured deliverable. It mentions a reference case to illustrate scope. However, it does not explicitly differentiate from sibling tools like sustainability_report or carbon_footprint_calculator, though the purpose is distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description mentions 'Gapup agent-payable C-suite expertise' hinting at a premium tool, but no when-not-to-use or comparison to siblings. The instruction to 'send the documented case fields' is about parameter input, not usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sabbatical_policy_comparatorA
Read-onlyIdempotent
Inspect

Enables CHROs to benchmark their company's sabbatical policies against peer organizations using data from SHRM, Payscale, and Mercer. Inputs include company size, industry, and current policy details. Outputs structured comparison with cost impact analysis, eligibility criteria, and duration benchmarks. Ideal for strategic HR planning and policy optimization.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
industryYesIndustry classification code (NAICS)
peerGroupNoList of peer company names for direct comparison
companySizeYesNumber of employees in the company
currentPolicyYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
benchmarkNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, covering safety and idempotency. The description adds value by detailing the output (cost impact analysis, eligibility criteria, duration benchmarks) and data sources, beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each adding value: purpose, data sources, inputs, outputs, and ideal use case. Front-loaded with the core function, no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, nested object, output schema exists), the description is complete. It covers purpose, inputs, outputs, and strategic context, and annotations cover behavioral traits. The output schema handles return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high (80%), with parameter descriptions already present. The description summarizes inputs (company size, industry, current policy details) but does not add new meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's specific action: benchmarking sabbatical policies against peers using data from SHRM, Payscale, and Mercer. It distinguishes from similar sibling tools like 'executive_comp_peer_benchmark' by focusing on sabbaticals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Ideal for strategic HR planning and policy optimization,' providing context for when to use. It does not explicitly state when not to use or list alternatives, but the niche focus implies appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

safety_guardrail_breach_analyzerA
Read-onlyIdempotent
Inspect

Analyzes potential LLM guardrail breaches against IEEE 7000 ethical compliance standards. Designed for risk persona to evaluate safety violations in AI outputs. Accepts raw LLM responses or structured breach reports, returns compliance analysis with severity scoring and mitigation recommendations.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
contextNoContextual information about the prompt or conversation
llmOutputYesRaw text output from LLM to analyze for guardrail breaches
severityThresholdNoMinimum severity score to report (0-10 scale)
includeMitigationsNoWhether to include mitigation recommendations

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
breachesNo
warningsNo
complianceScoreNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is clear. The description adds valuable context beyond annotations: it specifies the analysis standard (IEEE 7000) and output content (severity scoring, mitigation recommendations). There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of three concise sentences, front-loading the core purpose. Every sentence adds essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and full schema coverage, the description adequately covers purpose, inputs, outputs, and target user. It lacks details on error handling or edge cases, but these are not critical for a well-annotated tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, baseline is 3. The description adds meaning by clarifying that 'llmOutput' accepts both raw text and structured reports, and linking 'severityThreshold' and 'includeMitigations' to the output components. This exceeds what the schema alone provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Analyzes') and resources ('LLM guardrail breaches against IEEE 7000 ethical compliance standards'), clearly distinguishing it from siblings like 'safety_violation_incident_logger' (logs) and 'jailbreak_attempt_detector' (detects jailbreaks). It also specifies the output components ('compliance analysis with severity scoring and mitigation recommendations').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions the target persona ('risk persona') and input types ('raw LLM responses or structured breach reports'), but it does not explicitly state when to use this tool versus alternatives like 'ai_act_incident_response' or 'bias_amplification_tracker'. More precise exclusionary guidance would improve it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

safety_violation_incident_loggerC
Read-onlyIdempotent
Inspect

Logs AI safety violations for compliance reporting, targeting risk management personas. Accepts incident details such as violation type, severity, description, and timestamp. Returns structured data with compliance categorization based on NIST AI RMF guidelines. Ideal for automated incident tracking and regulatory reporting workflows.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
metadataNo
severityYes
timestampYes
descriptionYes
violationTypeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
incidentIdNo
nistReferenceNo
complianceCategoryNo
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states 'Logs AI safety violations,' implying a write operation, but annotations declare readOnlyHint=true, creating a direct contradiction. This severely undermines transparency. Beyond the contradiction, no behavioral traits like auth needs or side effects are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three sentences, front-loading the main purpose. It is efficient but could benefit from structured formatting (e.g., bullet points) for key details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema and mentions returning structured data with NIST categorization, which is helpful. However, the contradiction between description and annotations creates confusion about actual behavior. Given many safety-related siblings, more disambiguating context is needed for completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%, and the description merely lists the required parameters (violationType, severity, description, timestamp) without adding semantic details like format constraints or allowed values beyond what enums provide. It fails to compensate for low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool logs AI safety violations for compliance reporting and mentions targeting personas. It includes specific verb (logs) and resource (AI safety violations). However, it does not explicitly differentiate from similar sibling tools like bias_amplification_tracker or safety_guardrail_breach_analyzer, lacking direct sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes it is ideal for automated incident tracking and regulatory reporting workflows, providing some context. However, it offers no explicit guidance on when not to use it or alternatives to consider. The usage context is implied but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sales_enablement_architectC
Read-only
Inspect

Architecte Sales Enablement — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Spendesk — 45 reps · attainment 67% · ramp 5 mois → 3 mois · programme 8 modules · +€2,1M ARR. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
gapsYes
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
salesTeamYes
objectivesYes
currentEnablementYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side and returns a structured deliverable, but doesn't disclose rate limits, auth needs, or other behavioral traits. Adds marginal value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise with a reference case, but could be more structured. The reference case adds length without deep utility. Front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters with nested objects, no output schema), the description is insufficient. It doesn't describe the deliverable format, what the audit contains, or how to interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%. The description does not explain individual parameters or their meaning beyond the schema field names. It says 'send the documented case fields' but does not compensate for low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool returns a structured, audited sales enablement deliverable for C-suite expertise (CRO). It is specific about the resource and verb, but does not distinguish from sibling tools like revops_architect or abm_architect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for sales enablement architecture but provides no explicit guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sales_pipeline_forecastC
Read-only
Inspect

Prévision de pipeline commercial — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Doctolib Enterprise — pipeline Q2 2026 · 50 deals enterprise/mid-market · forecast confidence par deal + commit/best-case/worst-case. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
pipelineYes
historicalConversionByStageNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the description's mention of server-side validation and 'structured, audited deliverable' adds some context. However, it does not disclose whether the tool is slow (despite having an async parameter) or other behavioral nuances.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short but includes a specific reference case (Doctolib Enterprise) that may not generalize and adds unnecessary detail. It is not as concise as it could be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a complex nested input schema and no output schema. The description only vaguely mentions the output as a 'structured, audited deliverable' without specifying fields or structure, leaving significant gaps for an agent to understand what results to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 20%, the description should compensate by explaining parameter usage, but it only says 'send the documented case fields' without detailing individual parameters. This adds negligible value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool forecasts a sales pipeline and returns a structured, audited deliverable with confidence per deal and commit/best-case/worst-case. While it doesn't explicitly distinguish from siblings, there are no directly competing tool names that overlap precisely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lacks guidance on when to use this tool versus alternatives. It mentions a reference case but does not provide contextual requirements or conditions for ideal usage. No exclusions or comparisons to sibling tools are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sanctions_screener_multiA
Read-only
Inspect

Screening Sanctions Multi-listes — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Answers: For , run full OFAC + EU + UK HMT + UN + SECO + Canada SEMA + PEP + adverse media screening with composite risk score and evidence trail. · Is <company/individual> on any major international sanctions list? · What is the composite AML risk score for across all major watchlists? · Screen this M&A target / supplier / LP against all major sanctions lists and give me a compliance recommendation. · Is a PEP or associated with a PEP? What Enhanced Due Diligence is required? Reference case: Veridian Trading Co. LLC (Cyprus) — 7 listes · PEP check · adverse media 2 ans · composite 52/100 · escalate-to-compliance → EDD requis. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
addressNo
aliasesNo
entity_nameYes
entity_typeYes
context_noteNo
date_of_birthNo
jurisdiction_focusYesall
country_of_registrationNo
adverse_media_lookback_daysYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and open-world behavior. The description adds that it returns a 'structured, audited deliverable' with evidence trail, disclosing the output nature. No contradiction with annotations. The description enhances transparency beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately concise but includes a confusing first sentence ('Gapup agent-payable C-suite expertise (RISK)') and a lengthy reference case. It is structured with bullet points but could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and 10 parameters, the description provides use cases, list coverage, and a reference case. However, it lacks details on output format (only mentions 'structured, audited deliverable') and does not explain optional parameters like address or aliases. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 10% (only async has description). The description mentions required fields indirectly via examples but does not explain each parameter's meaning. It adds some value by listing example inputs and mentioning validation server-side, but insufficient for full parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool screens entities against multiple sanctions lists, PEP, and adverse media, returning a composite risk score. It lists specific lists (OFAC, EU, UK HMT, UN, SECO, Canada SEMA) and example use cases, differentiating it from siblings like kyc_screener.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides several example questions that indicate when to use (e.g., 'Is <company/individual> on any major sanctions list?', 'Screen this M&A target...'). It also mentions a reference case. However, it does not explicitly state when not to use or provide direct comparisons to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_playsC
Read-only
Inspect

Plans de sauvetage clients — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Kyriba — Plan sauvetage 30j · ARR €11.988 · Champion parti · Script 6 actions · 3 concessions. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
accountYes
companyYes
productYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint. The description adds that inputs are validated server-side and mentions a reference case, but does not explain async behavior or non-determinism implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but mixes French and English and includes a messy reference case. It conveys the core purpose without excessive length, but could be more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nested objects, async flag, and no output schema, the description lacks details on expected input structure, async usage, and deliverable contents, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (async has a description). The description does not elaborate on any parameters beyond 'send the documented case fields', leaving critical nested objects and required fields unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns a structured, audited deliverable for client rescue plans and provides a reference case. The purpose is clear, though 'save_plays' as a name is ambiguous and the output format is not fully detailed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like 'churn_defender' or 'renewal_optimizer'. Context is implied from the title and reference case but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sec_filing_decoderC
Read-only
Inspect

Décodeur de filing SEC — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Answers: Read the 10-K of and give me the material red flags, KPI movements, and a board-ready executive summary. · What has materially changed in 's risk profile in its latest annual filing? Flag any going-concern or auditor-change signals. · Is there any M&A signal or strategic review hint in 's most recent SEC filings? What's the evidence? · Prepare a due-diligence SEC filing brief for : financial snapshot, red flags, governance changes, and recommended next actions. · What is the sentiment of 's latest 10-K compared to its most recent 10-Q — bullish, neutral, or bearish? Reference case: SHOP · 10-K FY2024 · 4 red flags (1 critical: merchant concentration) · Revenue +24.7% YoY · . Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
cikNo
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusYesall
tickerNo
filing_typesYes
lookback_monthsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that it returns a 'structured, audited deliverable' and validates inputs server-side, which provides some behavioral context. However, it does not elaborate on rate limits, authentication, or what 'audited' entails. With annotations covering the basics, this is adequate but not excellent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively long and includes a list of example queries and a reference case. It is front-loaded with the purpose but lacks a clear structure (e.g., no separation of purpose, usage, parameters). Some sentences could be consolidated without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (SEC filing analysis with multiple focus areas), the description does not fully explain the output format or what 'structured, audited deliverable' means. It lacks details on how results are presented, whether they include numerical data or narrative, and how to interpret them. The absence of an output schema increases the need for such context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%, with only 'async' described. The description uses placeholders like <ticker> but does not explain each parameter's purpose or format. For example, 'focus' enum values are not elaborated. The description fails to compensate for the low schema coverage, leaving many parameters ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool decodes SEC filings, specifically focusing on red flags, KPI movements, and executive summaries. It provides example queries that illustrate its purpose. However, it does not explicitly differentiate from sibling tools like earnings_reviewer or competitive_deep_dive, which could perform similar analyses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes example questions but provides no explicit guidance on when to use this tool versus alternatives. It does not mention when not to use it or suggest other tools for related tasks. The agent would need to infer usage from examples, which is insufficient for optimal selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sentiment_news_pulseB
Read-only
Inspect

Pulse Média & Sentiment — Gapup agent-payable C-suite expertise (CMO). Returns a structured, audited deliverable. Answers: What is the current PR / brand sentiment for over the last 7 days? Show top headlines, trend signals, and recommended actions. · Is there a crisis building for ? Detect early-warning signals in press coverage and flag emerging negative narratives. · Track launch media coverage for — what is the press sentiment and which topics dominate the conversation? · Compare media sentiment between and its competitors over the past week. · What should our communications director prioritize in the next 48h based on current press coverage of ? Reference case: Velora Payments — Pulse média 7j · sentiment neutre (score +5) · crise émergente détectée · . Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
entity_nameYes
entity_typeYescompany
sentiment_lensYesreputation
date_range_daysYes
language_filterYesen
include_competitorsNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true. The description adds that it 'Returns a structured, audited deliverable' and mentions server-side validation. However, it does not clarify the async parameter behavior, rate limits, or what 'audited' entails. With annotations covering safety, the description adds moderate value but not rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat verbose and includes a marketing-like opening ('Gapup agent-payable C-suite expertise (CMO)') and a case reference. While the structure lists questions, it could be more concise and focused. The key functional statement ('Returns a structured, audited deliverable') is present but buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters, 5 required, no output schema, and low schema coverage, the description provides use cases and example output (score +5) but does not explain the full output structure or mention the async parameter behavior. The tool is moderately complex; the description fills some gaps but leaves important details (e.g., return format, async polling) unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% (only 'async' described). The description provides implied context for parameters like date_range_days (7 days), entity_name (company, brand), sentiment_lens (crisis-detection, launch-monitoring), and include_competitors (comparison). However, it does not explicitly map parameters to their schema definitions, leaving ambiguity. It adds meaning but falls short of systematic documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool provides media sentiment analysis for entities like companies, brands, and products, with specific use cases (e.g., crisis detection, launch tracking). It clearly identifies the resource (sentiment of news media) and the action (pulse/analysis). However, it does not explicitly differentiate from sibling tools like 'reputation_engine' or 'press_influencer', so it's slightly less than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists multiple example questions that imply when to use the tool (e.g., 'What is current PR sentiment?', 'Is there a crisis building?'). It provides context for typical use cases but no explicit guidance on when not to use it or which sibling tools to prefer for related tasks. The lack of exclusions or alternatives prevents a higher score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seo_cro_auditA
Read-only
Inspect

Full SEO + CRO audit of any public URL. Analyses technical SEO (HTTP status, HTTPS, title/meta/canonical/robots, H1-H2, JSON-LD structured data, sitemap, robots.txt, OG/Twitter cards), content SEO (word count, keyword density top-10, readability estimate, image alt coverage, internal/external links), performance signals (page size, estimated render time, inline scripts/styles, unoptimised images), and CRO (CTA detection, above-fold CTAs, forms, social proof, trust signals, pricing visibility). Optionally compares up to 5 competitor URLs. Returns 0-100 scores per dimension plus a prioritised (P0/P1/P2) recommendation list. ICP: marketing managers, SEO/CRO consultants, e-commerce ops, agency teams. Budget: 8s per URL. Cache TTL: 1h.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesFully-qualified URL to audit (e.g. https://stripe.com/pricing)
modeNoAudit scope — defaults to 'full'
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
compare_competitorsNoOptional list of competitor URLs to compare (max 5)

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYes
statusYes
sourcesYes
audit_modesYes
content_seoYes
cro_signalsYes
quality_scoreYes
technical_seoYes
overall_scoresYes
recommendationsYes
performance_signalsYes
competitor_comparisonNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false, and description aligns (no contradiction). Adds valuable behavioral details: async support, budget of 8s per URL, cache TTL of 1h, and return structure (0-100 scores with prioritised recommendations).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is front-loaded with main purpose and structured into clear categories. Each sentence adds value, though it is relatively long. Could be slightly trimmed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 params, output schema exists), the description is thorough: covers all audit dimensions, output format, operational constraints (budget, cache), target audience, and optional features. No major gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. Description provides context for the url parameter (what the audit covers) and mentions optional competitor comparison, but does not add significant new meaning beyond the schema's property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Full SEO + CRO audit of any public URL' with detailed breakdown of technical, content, performance, and CRO aspects. Distinguishes from sibling tools like seo_keyword_research and competitive_deep_dive by being a comprehensive audit covering both SEO and CRO.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly identifies target ICP (marketing managers, SEO/CRO consultants, e-commerce ops, agency teams) and optional competitor comparison. Provides clear context for use, though lacks explicit when-not-to-use statements or alternatives beyond the implicit differentiation from siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seo_keyword_researchA
Read-only
Inspect

SEO keyword research from a seed keyword or topic. Uses Google Suggest (public, keyless) to discover related queries at 2 expansion levels, then clusters them by intent: informational / commercial / transactional / navigational — via heuristic pattern matching. Search volume is bucketed (very_high / high / medium / low / very_low) and clearly labelled as ESTIMATED — no fabricated precise numbers. Returns all keywords, intent clusters, quality scores (0-100), and top 10 opportunities. Supports country (gl) and language (hl) targeting. 100% keyless. Cache TTL 6h. ICP: SEO managers, content strategists, SaaS founders, agency teams.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
countryNoISO 3166-1 alpha-2 country code for Google Suggest (e.g. 'US', 'FR', 'DE'). Defaults to 'US'.
languageNoBCP-47 language code for suggestions (e.g. 'en', 'fr', 'de', 'es'). Defaults to 'en'.
seed_keywordYesThe seed keyword or topic to research (e.g. 'invoice software', 'project management tool')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
countryYes
clustersYes
languageYes
warningsYes
all_keywordsYes
seed_keywordYes
quality_scoreYes
total_keywordsYes
top_opportunitiesYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond annotations by detailing the method (Google Suggest, heuristic clustering), volume bucketing (labelled ESTIMATED), cache TTL, and output specifics (intent clusters, quality scores). There is no contradiction with annotations (readOnlyHint=true, openWorldHint=true).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and concise despite its length. It front-loads the purpose, then method, output, constraints (keyless, caching), and ICP. Every sentence adds value with no repetition or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of annotations and output schema, the description is complete. It covers input parameters, underlying method, output specifics, caching behavior, and target audience. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds examples for 'seed_keyword' and explains the purpose of country/language parameters in the context of Google Suggest. It adds marginal value beyond the schema by contextualizing the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool does 'SEO keyword research from a seed keyword or topic' and explains the method (Google Suggest, intent clustering). The name and description together make the purpose immediately clear, and it distinguishes itself from sibling tools like 'seo_cro_audit' by being a research tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies the tool uses Google Suggest (public, keyless), supports country and language targeting, and targets ICP (SEO managers, etc.). It mentions cache TTL and that it's 100% keyless. However, it does not explicitly state when NOT to use it or provide alternatives for different keyword research needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sharia_compliance_screenerA
Read-only
Inspect

Sharia compliance screening engine for Islamic banks, Sukuk issuers, Gulf sovereign funds, halal investment managers and MENA family offices. Zero competing MCP on this vertical.

Standards supported: AAOIFI (default) | MSCI_Islamic | S&P_Sharia | DJIM

Four modes: • company — Full Sharia screen of a listed company: business activity (halal/haram/mixed) + AAOIFI financial ratios (debt/market-cap <30%, interest-assets <30%, non-compliant revenue <5%) • instrument — Sukuk / halal fund classification by ISIN or name. Maps to known Sharia boards. • sector_screen — Industry classification (halal/haram/mixed) with rationale + examples. Static AAOIFI-based map covering 40+ sectors. • financial_ratios — AAOIFI ratio computation on fetched or provided financials.

Prohibited activities screened: alcohol, gambling, pork, weapons, pornography, tobacco, conventional banking (riba), conventional insurance, adult entertainment, embryonic stem cells.

Output includes compliance_status (halal/haram/doubtful_mixed/purification_required), purification_pct when applicable, P0/P1/P2 signals, quality_score, and sources.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesScreening mode. company=full listed company screen, instrument=Sukuk/fund classification, sector_screen=industry halal/haram classification, financial_ratios=AAOIFI ratio check.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesEntity to screen. Company name, ticker or ISIN (e.g. "Aramco", "AAPL", "tobacco", "XS1234567890").
standardNoSharia standard to apply. Default "AAOIFI" (most conservative, widely accepted by Islamic banks).

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
companyNo
signalsYes
sourcesYes
instrumentNo
quality_scoreYes
sector_screenNo
standard_usedYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds behavioral details such as supported standards, prohibited activities, and output fields (compliance_status, purification_pct, etc.), which provide useful context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with bullet points for modes, standards, and prohibited activities. It front-loads purpose and is efficiently organized, though slightly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 modes, multiple standards, prohibitions) and the existence of an output schema (not shown but mentioned), the description fully explains the tool's capabilities, inputs, and outputs. No gaps are apparent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with all parameters described. The description adds extra meaning, such as default standard (AAOIFI) and why it is preferred, and the max length constraint. This adds value beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a 'Sharia compliance screening engine' with four well-defined modes. It explicitly claims 'Zero competing MCP on this vertical,' distinguishing itself from all sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each mode (company, instrument, sector_screen, financial_ratios) with examples. It does not explicitly state when not to use, but the specificity and the claim of no competitors implies clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

social_engagement_velocity_trackerA
Read-onlyIdempotent
Inspect

Tracks hourly social engagement velocity (likes, shares, comments) across Twitter, LinkedIn, and Reddit for CMOs. Inputs include platform handles/subreddits and time range. Outputs engagement metrics, velocity trends, and platform-specific insights. Ideal for real-time marketing performance monitoring and competitive benchmarking. Keywords: social media analytics, engagement tracking, marketing KPIs, CMO dashboard.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
hoursNo
platformsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
trendsNo
sourcesNo
warningsNo
engagementNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly and idempotent hints. The description adds context by specifying tracked metrics (likes, shares, comments) and outputs (velocity trends, platform insights), consistent with annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences with front-loaded action. The keywords line at the end is slightly repetitive but does not detract significantly. Could be slightly more focused avoiding keyword stuffing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema, the description adequately summarizes outputs as engagement metrics, velocity trends, and insights. For a 3-parameter tool, this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 33% with descriptions on 'async' only. The tool description mentions time range and platform handles/subreddits, adding some value beyond the schema but not fully compensating for low coverage. Parameter details remain basic.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it tracks hourly social engagement velocity across specific platforms (Twitter, LinkedIn, Reddit) for CMOs, with a specific verb and resource. It distinguishes itself from siblings by focusing on engagement metrics and target audience.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It indicates ideal use for real-time marketing monitoring and competitive benchmarking, providing context. However, it does not explicitly state when not to use or compare to alternative tools in the sibling list, such as brand_builder or sentiment_news_pulse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

social_influencer_fake_follower_detectorA
Read-onlyIdempotent
Inspect

Analyzes up to 10 social media influencers for fake followers by checking engagement velocity patterns (Trends24) and RSS feed anomalies. Returns authenticity scores, follower growth spikes, and suspicious activity flags. Optimized for CMOs evaluating influencer partnerships. Includes keywords: influencer marketing, fake follower detection, engagement analysis, social media audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
platformYesSocial media platform of the influencers
influencerHandlesYesArray of up to 10 social media handles (e.g., ['@influencer1', 'user2'])

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
resultsYes
sourcesYes
summaryNo
warningsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the description doesn't need to repeat them. It adds context about the analysis methods (Trends24, RSS anomalies) and outputs, but does not disclose limitations or edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with the main action, and includes only necessary details. Each sentence adds value, including the keywords section which can aid searchability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema, the description provides a good overview of what it returns. It is complete enough for a simple parameter set, though it could mention potential limitations or accuracy notes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents parameters adequately. The description does not add significant meaning beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes up to 10 influencers for fake followers using engagement velocity patterns and RSS anomalies, returning specific outputs. It distinguishes itself from sibling tools like 'social_engagement_velocity_tracker' by focusing on fake follower detection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions it's optimized for CMOs evaluating influencer partnerships, implying a use case, but does not explicitly state when to use it versus alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sovereign_data_breach_impactA
Read-onlyIdempotent
Inspect

Estimates financial impact of a data breach across three jurisdictions (US, EU, UK) for CFO strategic planning. Inputs include breach size, industry sector, and affected jurisdictions. Outputs include direct costs, regulatory fines, reputational damage, and cyber insurance premium adjustments. Ideal for cross-border risk assessment, financial contingency planning, and board-level reporting. Keywords: data breach cost, regulatory fines, cyber insurance, financial risk, cross-jurisdiction impact.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
industryNoIndustry sector of the affected organization
records_lostYesNumber of records compromised in the breach
jurisdictionsYesJurisdictions where the breach has legal or financial impact
detection_time_daysNoTime in days to detect the breach

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
total_cost_usdNoEstimated total financial impact in USD
cost_per_record_usdNoCost per compromised record in USD
regulatory_fines_usdNo
cyber_insurance_impactNo
reputational_damage_usdNoEstimated reputational damage cost in USD
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true, openWorldHint=true, and idempotentHint=true, signaling a safe, read-only, idempotent tool. The description adds valuable context about what the tool estimates and outputs (direct costs, regulatory fines, etc.), which goes beyond the annotations. It does not contradict any annotation, so a score of 4 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences plus keyword tags. It is front-loaded with the core function ('Estimates financial impact...') and efficiently summarizes inputs, outputs, and use cases. Every sentence adds value, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters (2 required), an output schema, and comprehensive annotations, the description is complete enough. It explains the purpose, required inputs, outputs, and ideal use cases. The presence of an output schema (not shown but referenced) means the description need not detail return values. A score of 4 reflects that it covers all essential aspects without gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mentions 'breach size, industry sector, and affected jurisdictions' as inputs, which maps to records_lost, industry, and jurisdictions parameters. However, it adds no additional meaning beyond what the schema already provides for each parameter. No improvement over baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'estimates financial impact of a data breach across three jurisdictions (US, EU, UK) for CFO strategic planning.' It specifies inputs (breach size, industry sector, jurisdictions) and outputs (direct costs, regulatory fines, etc.). This distinguishes it from sibling tools, none of which directly address data breach financial impact estimation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear use cases: 'cross-border risk assessment, financial contingency planning, and board-level reporting.' It does not explicitly state when not to use or name alternatives, but the specificity of the tool and sibling list make alternatives unnecessary. The context is well-defined, justifying a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sre_slo_breach_predictorA
Read-onlyIdempotent
Inspect

As a CTO, predict potential SLO breaches 24 hours in advance by analyzing public incident reports and MITRE ATT&CK techniques. Input your service's critical components and reliability thresholds to receive breach probability scores, top contributing TTPs, and recommended mitigations. Uses MITRE ATT&CK, GitHub Advisories, and Cloudflare Radar data. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
time_window_hoursNo
service_componentsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
incident_reportsNo
breach_probabilityNo
recommended_actionsNo
top_ttp_contributorsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint. Description adds value by specifying data sources (MITRE ATT&CK, GitHub Advisories, Cloudflare Radar) and async behavior, complementing annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose, no wasted words. Each sentence adds value: purpose, inputs/outputs, data sources, and async tip.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and presence of output schema, the description fully covers inputs, outputs, data sources, and async option. No additional context needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 33% (only async described). Description adds meaning for service_components ('critical components and reliability thresholds') and async, and implies time_window_hours through '24 hours in advance'. Compensates for low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it predicts SLO breaches 24 hours in advance using specific data sources and outputs scores, TTPs, and mitigations. It distinguishes from siblings by focusing on SLO breach prediction, which is unique among the listed tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides context for CTOs and advises using async to avoid timeout. Lacks explicit when-not-to-use or alternatives, but the tool's niche purpose makes usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

strategic_options_analyzerB
Read-only
Inspect

Analyseur d'options stratégiques — Gapup agent-payable C-suite expertise (CSO). Returns a structured, audited deliverable. Reference case: Aircall — 5 options stratégiques post-Série D (2023-2024). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
optionHypothesesYes
strategicContextYes
founderConstraintsYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side and that it returns a 'structured, audited deliverable'. It does not disclose response format, cost, or side effects beyond what annotations imply, but adds moderate context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, mixing French and English. It front-loads the purpose but includes unnecessary details (reference case). It could be more concise and better separated from example data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description should explain the return value beyond 'structured, audited deliverable'. This is insufficient for an agent to know what to expect. Additionally, the complex parameters are not explained, leaving gaps in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 20%, very low for a complex nested schema. The description does not explain individual parameters or their usage, saying only 'send the documented case fields'. This does not compensate for the lack of parameter documentation in the schema, hindering the agent's ability to construct correct inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes strategic options and returns a structured deliverable ('Analyseur d'options stratégiques', 'Returns a structured, audited deliverable'). It references a case study to illustrate capability. However, it does not differentiate from siblings beyond the 'CSO' mention; there are many similar strategy tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for C-suite strategic options analysis with 'Gapup agent-payable C-suite expertise (CSO)'. It provides context about server-side validation and a reference case. No explicit when-not-to-use or alternative tools are mentioned, leaving guidance minimal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

supplier_esg_auditC
Read-only
Inspect

Audit ESG des fournisseurs — Gapup agent-payable C-suite expertise (SUSTAINABILITY). Returns a structured, audited deliverable. Reference case: TechCorp — Audit ESG fournisseurs 2025 (5 fournisseurs, €1.37M spend). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
suppliersYes
targetScoreNo
auditCriteriaYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that inputs are validated server-side and returns an audited deliverable, which does not contradict annotations. No further behavioral details are provided beyond what annotations imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is short but includes a reference case that may add context. However, some phrasing is marketing-like ('Gapup agent-payable C-suite expertise'). Could be more concise and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested objects, no output schema, low schema coverage), the description is incomplete. It does not explain the return format, how to interpret the deliverable, or provide enough context for the agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (17%). The description does not compensate by explaining parameter meanings or usage. It only vaguely references 'documented case fields' without elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it audits suppliers for ESG and returns a structured deliverable. The verb 'Audit' and resource 'fournisseurs' are specific, but it does not differentiate from sibling tools like 'esg_audit_multi'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. The description only mentions server-side validation and a reference case, but does not specify prerequisites or when to avoid using it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

supply_chain_fx_exposure_dashboardB
Read-onlyIdempotent
Inspect

Provides real-time foreign exchange exposure dashboard for supply chain monitoring. Designed for COO persona to track currency risk across suppliers and regions. Inputs include supplier IDs, base currency, and target currencies. Outputs structured FX exposure data with risk indicators, exchange rates, and supplier impact analysis sourced from World Bank LPI and live FX rate APIs.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
supplierIdsNoList of supplier identifiers to analyze
baseCurrencyYesBase currency code (ISO 4217) for exposure calculation
riskThresholdNoPercentage threshold for high-risk exposure flagging
targetCurrenciesYesTarget currency codes (ISO 4217) to compare against base

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
statusYes
sourcesNo
warningsNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, idempotentHint, and openWorldHint. The description adds that the tool uses real-time data from World Bank LPI and live FX APIs, which is useful context. However, no further behavioral traits (e.g., rate limits, caching) are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no wasted words. Front-loaded purpose and persona description. Could be more structured with bullet points but still concise and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool complexity (5 parameters, many siblings), the description covers core purpose and data sources but omits mentioning the riskThreshold and async parameters, which are part of the schema. Output schema exists so return values are covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions cover 100% of parameters, so baseline is 3. The description repeats parameter info (supplier IDs, base currency, target currencies) but does not add new meaning beyond the schema. It groups them but no additional semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it provides an FX exposure dashboard for supply chain monitoring, listing inputs and outputs. However, it does not explicitly distinguish from sibling tools like fx_rate or working_capital_fx_hedge_optimizer beyond the dashboard focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description identifies the COO persona but provides no explicit guidance on when to use this tool versus alternatives. No 'use when' or 'avoid when' criteria, and no mention of prerequisites or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sustainability_reportC
Read-only
Inspect

Rapport de durabilité — Gapup agent-payable C-suite expertise (SUSTAINABILITY). Returns a structured, audited deliverable. Reference case: GreenLoop Solutions — rapport durabilité B-Corp 2025 (95 FTE, €18M CA). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
pillarsYes
stakeholdersYes
targetLabelsNo
existingLabelsNo
audienceProfileYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description does not disclose key behaviors beyond what annotations provide. The presence of an async parameter is not mentioned in the text, and the description does not explain whether the operation is synchronous or polling-based. Annotations provide readOnlyHint and openWorldHint but no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short (3 sentences) but includes fluff like 'Gapup agent-payable C-suite expertise' which does not add value. It could be more concise and front-loaded with essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high complexity (8 parameters, nested objects, no output schema), the description is severely incomplete. It does not explain the output format, how to use the async option, or what the deliverable contains. The reference case provides some context but is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 13% schema description coverage, the description adds almost no meaning to the parameters. It merely mentions required fields ('company, pillars, stakeholders, audienceProfile') without explaining their role. The complex nested structure is not clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns a structured, audited sustainability deliverable and provides a reference case. However, it fails to differentiate from the sibling tool 'sustainability_reporting_pilot' which appears similar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. The description mentions 'Gapup agent-payable C-suite expertise' which is vague, and there is no mention of exclusions or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sustainability_reporting_pilotC
Read-only
Inspect

Pilote de reporting durabilité — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: AlphaTech Industries SAS — premier rapport CSRD wave 2 (exercice 2025). Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNo
companyYes
dataInputsYes
materialityYes
targetFrameworksYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations state readOnlyHint=true, but description implies report generation (potential side effect). No additional behavioral traits (auth, rate limits, etc.) are disclosed beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with some waste (jargon like 'Gapup agent-payable'). It is concise but lacks clarity; front-loading is moderate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complex tool (nested objects, 6 params, no output schema). Description is too brief; it omits deliverable structure, result handling, and relationships with siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 17% (very low). Description does not explain any parameters; it merely instructs to 'send the documented case fields' without adding meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states it is a sustainability reporting pilot that returns a structured audited deliverable, with a reference case. However, it lacks a clear verb and uses jargon ('Gapup agent-payable C-suite expertise (RISK)'); it does not distinguish from siblings like 'sustainability_report' or 'esg_audit_multi'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. No when-not or contextual recommendations are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syndicated_loan_covenant_breach_alertA
Read-onlyIdempotent
Inspect

Monitors syndicated loan covenants for potential breaches by analyzing Tradeweb market data. Designed for CFOs to proactively identify financial compliance risks in loan agreements. Accepts loan identifiers, covenant thresholds, and reporting period as inputs. Returns structured breach alerts with market context and severity indicators.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
loanIdYesUnique identifier for the syndicated loan
currencyNoISO currency code for financial values
reportingPeriodYesTime period for covenant compliance check
covenantThresholdsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
breachesNo
warningsNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, making the description's safety profile redundant. The description adds value by describing the output format (structured breach alerts with market context and severity indicators) but does not disclose additional traits beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long, front-loaded with the primary action, and every sentence serves a purpose (purpose, audience, inputs, outputs). No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (5 params, nested object, output schema), the description covers purpose, inputs, and output format. It does not explain the async parameter or return values in detail, but the output schema exists and annotations fill gaps, making it nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high (80%), so the schema already documents most parameters. The description mentions three required inputs (loan identifiers, covenant thresholds, reporting period) but does not add meaning beyond the schema. The baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('monitors') and resource ('syndicated loan covenants'), clearly stating the tool's purpose: identify potential breaches by analyzing Tradeweb market data. It distinguishes from sibling tools like bond_covenant_monitor by focusing on syndicated loans.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description targets CFOs for proactive compliance risk identification, providing some context. However, it lacks explicit when-to-use or when-not-to-use guidance and does not mention alternatives, leaving usage scope implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syndicated_loan_pricing_benchmarkA
Read-onlyIdempotent
Inspect

Provides CFOs with peer benchmarking for syndicated loan pricing by comparing current loan terms against market data from Tradeweb and FRED. Inputs include loan amount, tenor, credit rating, and currency. Outputs structured pricing benchmarks with spread, yield, and fee comparisons. Ideal for quick validation of loan competitiveness or negotiation preparation.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
tenorYesLoan tenor (e.g., '5Y', '3Y')
regionNoRegion for benchmarking (e.g., 'US', 'EU')
currencyYesCurrency code (e.g., 'USD', 'EUR')
loanAmountYesLoan amount in millions
creditRatingYesBorrower credit rating (e.g., 'BBB', 'BB+')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
benchmarksNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, idempotentHint, openWorldHint. Description adds context (data sources, outputs) but no behavioral traits beyond what annotations cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no wasted words, front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With output schema present, description sufficiently covers purpose, inputs, and outputs. Annotations confirm safety and idempotency, making it complete for this tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so description adds no additional parameter-level detail beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it provides peer benchmarking for syndicated loan pricing, specifying data sources (Tradeweb and FRED), inputs, and outputs. It distinguishes from sibling tools like syndicated_loan_covenant_breach_alert.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Mentions ideal use case for quick validation or negotiation preparation, but does not explicitly state when not to use or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

talent_contract_risk_mapperA
Read-onlyIdempotent
Inspect

For CHROs: analyzes employee contracts for non-compete, IP assignment, and confidentiality clauses, comparing against state labor laws and jurisdiction-specific precedents. Returns risk levels, conflicting statutes, and suggested revisions. Uses USPTO PatFT, CourtListener, and EUR-Lex for legal cross-referencing. Ideal for contract reviews, compliance audits, or policy updates.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
jurisdictionYesState or country jurisdiction (e.g., 'California', 'Germany')
contract_textYesFull text of the employee contract or clause section to analyze
employee_roleNoJob title or role classification (e.g., 'Software Engineer', 'Executive')
effective_dateNoContract effective date (YYYY-MM-DD)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
risk_summaryNo
suggested_revisionsNo
conflicting_statutesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint. The description adds that it uses external legal databases (USPTO, CourtListener, EUR-Lex) and returns risk levels and revisions, which complements the annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each adding distinct value: audience and core function, outputs, data sources and use cases. It is front-loaded and concise with no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 params, output schema exists, many siblings), the description covers the tool's purpose, outputs, sources, and use cases. It lacks explicit parameter guidance but that is covered by schema. It is sufficiently complete for selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are documented. The description does not elaborate on individual parameters but provides overall context. The added value is limited as the schema already explains parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes employee contracts for non-compete, IP assignment, and confidentiality clauses, comparing against jurisdictional laws. It specifies the return of risk levels, conflicting statutes, and suggested revisions, distinguishing it from sibling tools like contract_risk_scanner or legal_clause_extractor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates ideal use cases (contract reviews, compliance audits, policy updates) but does not explicitly exclude alternatives or contrast with siblings. It provides sufficient context for when to use, but lacks direct exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

talent_intelligenceA
Read-only
Inspect

HR tech intelligence for CHROs, recruiters, VC teams, comp & benefits leads and workforce planners. Four modes powered by ESCO, O*NET, BLS OES and crowd-sourced salary data:

• salary_benchmark — cash-only salary medians (p25/median/p75) for 54+ roles across US/EU/Asia. Covers tech, finance, compliance, healthcare, marketing, ops and C-suite. Data from BLS OES, Levels.fyi and StackOverflow Developer Survey 2024. • skills_taxonomy — maps a skill to its ESCO URI, O*NET codes, skill type (hard/soft/knowledge/cert), 8 related skills with similarity scores and typical roles. • job_market_trends — YoY growth %, open positions estimate, top employers and leading skills per job category × country. Static 2024 data with BLS baseline fallback. • adjacent_roles — up to 6 roles adjacent to a source role with ESCO taxonomy adjacency: similarity score, salary delta % and skills overlap %.

All salary data is cash-only (excludes equity/RSU/bonus). Cache TTL: 24h (stable labour market data). Optional env ONET_API_KEY for authenticated O*NET lookups (free registration at onetcenter.org).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesAnalysis mode: salary_benchmark=compensation data, skills_taxonomy=ESCO/O*NET mapping, job_market_trends=market growth and demand, adjacent_roles=career path recommendations.
roleNoJob title (required for salary_benchmark, job_market_trends, adjacent_roles). Examples: "Senior Software Engineer", "Compliance Officer", "Data Scientist", "CFO".
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
skillNoSkill to classify (required for skills_taxonomy mode). Examples: "Python", "transformer architecture", "GDPR", "Kubernetes", "leadership".
countryNoISO 2-letter country code. Default: US. Examples: US, FR, DE, GB, SG.
seniorityNoSeniority level. Default: senior. Affects salary benchmark ranges.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
sourcesYes
quality_scoreYes
adjacent_rolesNo
skills_taxonomyNo
salary_benchmarkNo
job_market_trendsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds behavioral details: data is static 2024 with BLS fallback, cache TTL 24h, optional ONET_API_KEY, and that salary data excludes equity/RSU/bonus. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points for each mode, making it scannable. It front-loads the purpose and target audience. While somewhat long, every sentence adds value and there is no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (four modes, multiple parameters, and an output schema), the description covers each mode, data sources, constraints (cash-only, cache TTL), and includes examples. It provides sufficient context for an agent to understand the tool's capabilities without gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds meaningful context for each parameter: role examples, skill examples, and ISO country code format. The mode enumeration is expanded with detailed descriptions of each mode's purpose and data sources, aiding the agent in selecting the correct mode.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides HR tech intelligence with four distinct modes, each with specific use cases. It distinguishes itself from siblings by naming concrete data sources (ESCO, O*NET, BLS OES, Levels.fyi) and target audiences (CHROs, recruiters, VC teams, etc.).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each mode with detailed examples (e.g., 'salary_benchmark — cash-only salary medians...'). It provides context for selecting the appropriate mode but does not explicitly say when NOT to use or compare to alternative tools among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

talent_litigation_exposureA
Read-onlyIdempotent
Inspect

Estimates litigation exposure risk for CHROs by analyzing past employee lawsuits, settlement amounts, and industry benchmarks. Inputs include company location, industry code, and employee count range. Returns exposure score, average settlement amounts, lawsuit frequency trends, and risk factors. Ideal for legal risk assessment, HR strategy planning, and board-level reporting. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
industry_codeYesNAICS industry code (e.g., '541511' for IT services)
employee_countNoCurrent number of employees
lookback_yearsNoNumber of years to analyze
company_locationYesState or region where company operates (e.g., 'CA', 'New York')

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
warningsYes
avg_settlementNoAverage settlement amount in USD
exposure_scoreYesNormalized risk score (0-100)
historical_trendNo
top_risk_factorsNo
lawsuit_frequencyNoLawsuits per 1000 employees per year
industry_benchmarkNoIndustry average exposure score
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true. The description adds value by noting the async option to avoid timeout, indicating potential slowness. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and front-loaded: first sentence states purpose, then inputs, outputs, use cases, and a practical tip. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 parameters, output schema, annotations), the description fully covers purpose, inputs, outputs, usage guidance, and a timeout avoidance tip. It is complete for effective tool selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description mentions key inputs but does not elaborate on all parameters (e.g., lookback_years). The async tip adds some value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (Estimates) and resource (litigation exposure risk), clearly distinguishing it from sibling tools like talent_contract_risk_mapper. It states the tool's function, inputs, outputs, and ideal use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context by listing ideal use cases (legal risk assessment, HR strategy planning, board-level reporting) but does not explicitly state when not to use it or mention alternatives among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

talent_poaching_riskA
Read-onlyIdempotent
Inspect

Analyzes employee poaching risk for CHROs by evaluating LinkedIn profile activity (job searches, profile views) and comparing compensation against BLS benchmarks. Returns a ranked list of high-risk employees with risk scores and suggested retention actions. Ideal for proactive talent retention strategies. Keywords: employee retention, poaching risk, compensation benchmark, LinkedIn activity, CHRO analytics.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
locationNoGeographic location filter (e.g., 'San Francisco, CA')
departmentYesDepartment filter (e.g., 'Engineering', 'Sales')
min_tenure_monthsNoMinimum tenure in months to include in analysis
benchmark_job_titleNoSpecific job title for compensation benchmarking

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
risk_assessmentNo
department_avg_riskNo
benchmark_comparisonNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, openWorldHint, idempotentHint) already indicate the tool is safe, idempotent, and uses external data. The description adds valuable context about analyzing LinkedIn activity and BLS benchmarks, enhancing transparency without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences plus keywords, front-loading the core purpose and output. Every sentence adds value with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema and moderate complexity. The description covers what it returns (ranked list with scores and actions) and its data sources, leaving minimal gaps. It could note prerequisites like LinkedIn data access, but overall it is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 5 parameters have descriptions in the input schema (100% coverage), so the schema already documents parameter meanings. The tool description does not add further semantic detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool analyzes employee poaching risk using LinkedIn profile activity and compensation benchmarks, returning a ranked list with risk scores and retention actions. It distinguishes itself from similar talent tools like talent_intelligence by its specific focus on poaching risk and proactive retention strategies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates it is ideal for proactive talent retention strategies and targets CHROs, providing clear usage context. However, it does not explicitly state when not to use this tool or mention alternative tools for different talent scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tariff_arbitrage_finderA
Read-onlyIdempotent
Inspect

As a COO, identify tariff reclassification opportunities to reduce import costs. Analyzes product HS codes against WTO TFA and USA Trade Online data to find lower-duty classifications. Inputs: product description, current HS code, country of origin, and annual import volume. Outputs: potential duty savings, alternative HS codes, and compliance considerations.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
annualVolumeNo
currentHsCodeYes
countryOfOriginYes
currentDutyRateNo
productDescriptionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
opportunitiesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, which the description aligns with (analyzing data, no side effects). The description adds context about data sources and output types, but does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences, with the first sentence front-loading the main purpose. Every sentence adds value: purpose, data sources, inputs/outputs. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity and low schema coverage, the description adequately explains inputs and outputs (duty savings, alternative codes, compliance considerations). It provides enough context for an agent to select and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at only 17%, the description compensates by listing four key parameters (productDescription, currentHsCode, countryOfOrigin, annualVolume) and their roles, adding meaning beyond the schema for most parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's primary purpose: 'identify tariff reclassification opportunities to reduce import costs.' It specifies the verb (identify), resource (tariff reclassification opportunities), and data sources (WTO TFA, USA Trade Online). The title and name distinguish it from siblings like tariff_impact_simulator and trade_finance_eligibility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides input and output lists, implying usage context ('As a COO'), but lacks explicit guidance on when to use this tool versus alternatives or when not to use it. No explicit when/when-not statements are present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tariff_impact_simulatorA
Read-onlyIdempotent
Inspect

As a COO, model how proposed tariff changes affect landed costs for imported goods. Inputs: HS code, current tariff rate, proposed tariff rate, product value, shipping cost, and country of origin. Outputs: detailed cost breakdown including new duties, taxes, and total landed cost impact. Sources include WTO TFA and US Census trade data.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
hsCodeYes
productValueYes
shippingCostNo
countryOfOriginYes
currentTariffRateYes
proposedTariffRateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
costImpactNo
currentDutyNo
proposedDutyNo
dutyDifferenceNo
currentLandedCostNo
proposedLandedCostNo
costImpactPercentageNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the safety profile is clear. The description adds value by stating the tool models cost impacts and cites sources (WTO TFA and US Census trade data), providing context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no fluff. It front-loads the primary purpose and efficiently covers inputs and outputs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, vague schema, output schema exists), the description adequately summarizes the main functionality. It covers essential inputs and outputs but omits details on the optional async parameter and explicit limitations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (14%), but the description lists all key inputs (HS code, tariff rates, product value, shipping cost, country of origin) and explains their role in the simulation. This compensates for the sparse schema and adds meaningful context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'model how proposed tariff changes affect landed costs for imported goods.' It specifies verb and resource, and lists inputs and outputs. However, it does not explicitly differentiate from sibling tools like 'tariff_arbitrage_finder', which could be confused as similar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It mentions the intended user role ('As a COO') but does not specify prerequisites, exclusions, or compare to other tariff-related tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tax_compliance_multiA
Read-onlyIdempotent
Inspect

Multi-jurisdiction tax compliance data for international SaaS, cross-border marketplaces and expat services. Five modes: (1) vat_lookup — validate EU VAT numbers live via VIES SOAP (27 EU countries) or UK VRN via HMRC; (2) sales_tax — US state sales tax rates, nexus thresholds (post-Wayfair 2018), digital goods taxability for all 50 states + DC; (3) gst — APAC GST/SST/consumption-tax rates for IN, SG, AU, NZ, MY, JP, KR, TH, ID, PH, VN with reduced rates and registration thresholds; (4) oss_ioss_eligibility — EU One-Stop-Shop and Import-OSS eligibility analysis (EUR 10k OSS threshold, EUR 150 IOSS per-consignment); (5) transfer_pricing_benchmark — OECD/JTPF operating-margin benchmarks by industry and country (20+ sectors, country-specific adjustments). Returns P0/P1/P2 compliance signals: P0=invalid VAT used for zero-rating, P1=taxable digital goods detected/audit risk, P2=filing deadlines/nexus alerts. Keyless — no API key required. Optional env: HMRC_VAT_API_KEY for UK VAT live validation. Cache TTL 24h.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesTax mode to invoke.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesMode-specific query: vat_lookup -> VAT number with country prefix (e.g. 'FR40303265045'); sales_tax -> US state code or name (e.g. 'CA', 'California'); gst -> ISO country code (e.g. 'SG', 'IN', 'AU'); oss_ioss_eligibility -> annual EU B2C revenue in EUR or keyword (e.g. '5000', 'below'); transfer_pricing_benchmark -> industry name (e.g. 'manufacturing', 'saas', 'r&d').
countryNoISO 3166-1 alpha-2 country code. Required for gst when query is ambiguous. Used in transfer_pricing_benchmark for country-specific OECD adjustments.
transaction_typeNoTransaction type for signal generation. 'digital' triggers GST/sales-tax digital goods warnings.

Output Schema

ParametersJSON Schema
NameRequiredDescription
gstNo
modeYes
statusYes
signalsYes
sourcesYes
oss_iossNo
sales_taxNo
vat_lookupNo
quality_scoreYes
transfer_pricingNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and idempotentHint=true, so no contradiction. The description adds valuable behavioral context: keyless access (no API key required), optional HMRC_VAT_API_KEY for UK VAT, 24h cache TTL, and the P0/P1/P2 signal taxonomy. It does not contradict annotations and enriches understanding beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with numbered modes and clear separation of concerns using bullet points. While lengthy, it earns its length by providing necessary detail for a complex multi-mode tool. Could be slightly more concise without losing clarity, but it's not verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 modes, 5 parameters, output schema present), the description covers all necessary aspects: mode selection, parameter semantics, behavioral traits (keyless, cache), return signals (P0/P1/P2), and optional configuration. With an output schema, the description need not explain return values, and it is fully adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by providing concrete examples for each mode's query parameter (e.g., 'FR40303265045', 'CA', 'manufacturing') and clarifying the purpose of the country and transaction_type parameters. This goes beyond the schema's minimal descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it provides multi-jurisdiction tax compliance data with five specific modes (vat_lookup, sales_tax, gst, oss_ioss_eligibility, transfer_pricing_benchmark). Each mode is distinctly described with its scope, making the purpose unambiguous and differentiating it from sibling tools, none of which are tax-related.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each mode, including example queries and specific contexts (e.g., EU VIES SOAP, US state sales tax). However, it does not explicitly state when not to use the tool or list alternatives, but given the uniqueness of the tool and clear mode descriptions, this is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tax_optimizationC
Read-only
Inspect

Optimisation fiscale — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: Pennylane — Fiscalité optimisée · CIR €1.2M · IP Box France 10% · Économie totale €2.4M/an. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
ipAssetsNo
activitiesYes
financialsYes
jurisdictionsYes
currentTaxOptimizationsNo
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. The description adds that it returns a structured, audited deliverable and inputs are validated server-side, but does not disclose other behavioral traits (e.g., rate limits, authentication needs, what happens on failure). The openWorldHint is left unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but contains jargon ('Gapup agent-payable C-suite expertise (CFO)') and an example case that may not be universally helpful. It is somewhat front-loaded with 'Optimisation fiscale' but could be clearer and more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested objects, 7 parameters, no output schema), the description is incomplete. It does not explain the deliverable's format, prerequisites, or how to interpret the result, leaving significant gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low (14%, only async is described). The description does not add meaning to the multiple nested parameters (company, financials, jurisdictions, etc.) beyond instructing to 'send the documented case fields,' which is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the tool as tax optimization and mentions a deliverable, but it uses vague French jargon ('Optimisation fiscale') and fails to clearly state the specific verb and resource. It does not distinguish itself from many tax-related siblings like tax_compliance_multi or ma_tax_efficiency_mapper.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description only notes that inputs are validated server-side, but lacks context on scenarios, prerequisites, or when not to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

term_sheet_negotiationC
Read-only
Inspect

Négociation term sheet — Gapup agent-payable C-suite expertise (FUNDRAISING). Returns a structured, audited deliverable. Reference case: Agicap Série C €50M — 8 clauses analysées · 3 rouges · Score fondateur 62/100 → plan pour 81. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
roundYes
companyYes
termSheetClausesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true. The description states it returns a deliverable, implying no side effects, but does not elaborate on behavioral traits like permissions, rate limits, or what happens to data. It is consistent but adds little beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences plus a reference example) and front-loaded with the purpose. However, the reference case example adds some length but not structure. Overall, it is efficient but could be more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested objects, no output schema), the description is inadequate. It does not explain the deliverable's structure, how to interpret the founder score, or provide example clause formats. The reference case hints but does not fully prepare the agent for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (25%), with only the 'async' parameter described. The description mentions 'send the documented case fields' but does not explain the nested objects (round, company, termSheetClauses) or their sub-fields. This forces the agent to infer meaning from names alone, which is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool negotiates term sheets for fundraising and returns a structured deliverable. It references a specific case (Agicap Série C) which adds context. However, the exact output format and scope are not fully specified, preventing a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The sibling list includes many fundraising-related tools (e.g., deal_coach, cap_table_strategist), but the description does not differentiate or provide selection criteria. It only mentions that inputs are validated server-side.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tool_recommendA
Read-only
Inspect

Cross-tool recommendation system: given a free-text intent, returns the most appropriate tools from the 170+ Gapup MCP catalogue, ranked by confidence, with pre-filled input suggestions and an optimal multi-tool chain when applicable. Use this first when you are unsure which tool to call — it navigates the full catalogue for you. Supports 15+ static pre-designed chains for frequent intents (M&A due diligence, sanctions screening, ESG 360, AI Act compliance, FTO patent clearance, crypto wallet tracking, etc.). Domains: compliance | finance | intel | legal | content | data | trade | infra. Pure compute — $0.01/call, no external fetch. Ideal as a first call in any multi-step agent workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
langNoOptional ISO 639-1 language hint (fr, en, de, zh, es …). Used for language-aware boosting.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
domainNoOptional domain hint to boost tools in this category.
intentYesFree-text description of what you want to accomplish. E.g. 'Run a full M&A due diligence on Acme Corp' or 'Je veux vérifier qu'un fournisseur n'est pas sous sanctions OFAC'. FR/EN/DE/ZH supported.
max_resultsNoMax number of recommendations returned (1-10). Default 5.
include_chainNoWhether to include a suggested_chain of tools in the optimal sequence. Default true. Chain is always included for well-known intents (M&A, compliance, ESG, etc.).

Output Schema

ParametersJSON Schema
NameRequiredDescription
intentYes
statusYes
sourcesNo
not_coveredNo
quality_scoreYes
recommendationsYes
suggested_chainNo
alternative_pathsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds value beyond the readOnlyHint annotation by detailing the tool's behavior: pure compute, $0.01/call, no external fetch, support for async and chains. It does not contradict annotations. Some details about response structure could be added, but the description is sufficiently transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the main purpose and uses bullet points for domains. It is informative but somewhat verbose; however, every sentence serves a purpose. Minor redundancy could be removed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 params, output schema exists), the description covers all key aspects: purpose, usage, cost, domains, chain support, async behavior. It is complete for an agent to understand when and how to use the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description reinforces high-level semantics (e.g., 'intent' is free-text, 'domain' is an enum) but does not add significant new meaning beyond what the input schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: a cross-tool recommendation system that returns the most appropriate tools based on a free-text intent. It uses specific verbs ('recommend', 'navigates') and distinguishes itself from siblings by being the first call when unsure which tool to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use: 'Use this first when you are unsure which tool to call' and positions it as ideal for multi-step workflows. It provides domain hints and chain support. However, it does not explicitly mention when not to use or provide specific alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trade_finance_eligibilityA
Read-onlyIdempotent
Inspect

Evaluates trade finance eligibility for CFOs by analyzing counterparty risk and jurisdiction using World Bank and BIS data. Inputs include counterparty country code (ISO 3166-1 alpha-3) and industry sector. Returns risk scores, eligibility flags, and financing terms. Ideal for assessing letters of credit, export credit agency guarantees, and other trade finance instruments. Keywords: trade finance, counterparty risk, jurisdiction risk, letters of credit, ECA guarantees.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
industrySectorYes
annualTradeVolumeUSDNo
counterpartyCountryCodeYes
counterpartyCreditRatingNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
eligibilityNo
financingTermsNo
countryRiskScoreNo
maxFinancingAmountUSDNo
recommendedInstrumentsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, so safety and idempotency are covered. The description adds that it returns 'risk scores, eligibility flags, and financing terms,' giving useful output context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences followed by keywords, which is reasonably concise. It front-loads the main action and audience, though the keywords list adds minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters (2 required) and an output schema exists, the description provides a good overview of purpose and output type. However, it omits descriptions for 3 optional parameters, which is a gap in completeness. The output schema partially fills return value understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only the 'async' parameter is described in schema). The description names 'counterparty country code (ISO 3166-1 alpha-3)' and 'industry sector' but fails to explain 'annualTradeVolumeUSD', 'counterpartyCreditRating', or the 'async' parameter. With low schema coverage, the description should compensate by detailing all parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it evaluates trade finance eligibility for CFOs using specific data sources (World Bank, BIS). It lists inputs and outputs and distinguishes itself from numerous sibling tools by focusing on trade finance instruments like letters of credit and ECA guarantees.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description explicitly says 'Ideal for assessing letters of credit, export credit agency guarantees, and other trade finance instruments,' providing clear ideal use cases. However, it does not mention when not to use or suggest alternative tools for similar but different needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribe_chapterize_mediaA
Read-onlyIdempotent
Inspect

Transcription and chapterization of long-form media (YouTube, podcasts, direct audio/video) for content marketing teams, podcast publishers, edu tech, journalists and accessibility/compliance.

Pipeline: • YouTube → timedtext captions (keyless) + oEmbed metadata + native timecode chapters from description • Podcast RSS → episode description + duration + timecodes if embedded in show notes • Direct media → partial (requires Whisper API via OPENAI_API_KEY + force_whisper:true) • Chapters: native YouTube timecodes preferred; heuristic TF-IDF segmentation as fallback • Summary: extractive TF-IDF top-sentences (no LLM required) • Language detection: character-set heuristic (CJK→zh, kana→ja, hangul→ko, accents→fr/de/es)

Output formats: json (full structured object) | text (plain transcript) | srt | vtt

SLA: ≤15s budget total. Cache: 24h TTL.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube URL, podcast RSS feed URL, or direct MP3/MP4 URL. Example: "https://www.youtube.com/watch?v=jNQXAC9IVRw"
langNoISO 639-1 language hint (e.g. "en", "fr", "de"). Default "auto".
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
chapters_maxNoMaximum number of chapters. Default 8.
output_formatNoTranscript format. Default "json".
include_summaryNoInclude extractive summary. Default true.

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYes
statusYes
signalsYes
sourcesYes
summaryNo
chaptersYes
segmentsYes
key_topicsYes
transcriptYes
source_typeYes
lang_detectedYes
quality_scoreYes
duration_secondsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds significant behavioral context: pipeline steps (YouTube keyless captions, podcast oEmbed, direct Whisper API), chapterization fallback, summary method, language detection heuristics, output formats, SLA, and caching. This exceeds what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed and well-structured with bullet points for the pipeline steps. It is informative without being overly verbose, though some sentences could be trimmed for further conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, multiple sources, and an output schema), the description covers all essential aspects: sources, pipeline, parameters, SLA, caching, and output formats. The output schema exists, so return values do not need to be described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% description coverage, so baseline is 3. The description adds value by explaining how parameters like async, chapters_max, output_format, and include_summary affect the processing pipeline, providing context beyond the schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs transcription and chapterization of long-form media from YouTube, podcasts, and direct audio/video. It lists specific supported sources and the pipeline, distinguishing it from siblings by its scope and capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides guidance on when to use the tool for different media sources (YouTube, podcast RSS, direct media) and explains the chapterization preference (native timecodes first, then heuristic). It mentions SLA and caching but does not explicitly state when not to use the tool or list alternatives among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

treasury_optimizerC
Read-only
Inspect

Optimiseur de trésorerie — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: Alan — Trésorerie €380M post-Série F · Allocation optimale 4 instruments · Yield +145bp · +€5.5M/an. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
horizonNo
constraintsYes
cashPositionYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only (readOnlyHint=true) and open world (openWorldHint=true). The description adds minor context (server-side validation, structured deliverable) but does not disclose return format, authentication needs, or potential side effects beyond what annotations imply. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is adequately sized but includes a marketing reference case that adds little structural value. The first sentence identifies the tool, but the reference case and brand language reduce conciseness. Could be more direct.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 parameters with nested objects, no output schema), the description is incomplete. It fails to describe the deliverable's structure, how to interpret results, or any algorithmic context. The annotations provide some safety context but do not compensate for the lack of output schema or usage details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, but the description does not explain any parameters or their meaning beyond the schema's minimal descriptions. It only says 'send the documented case fields' without adding semantics, leaving the agent to rely solely on parameter names and types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a treasury optimizer for CFO-level expertise, returning a structured audited deliverable. It distinguishes itself from sibling tools by specifying the output type and including a reference case, though it does not explicitly differentiate from similar financial tools like working_capital.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool over alternatives. The description implies use for treasury optimization by CFOs but provides no exclusions or comparisons to sibling tools, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trend_watcherA
Read-onlyIdempotent
Inspect

Monitor emerging trends, regulatory shifts and adoption signals for a given market sector. Returns 5-12 trend cards, each with a momentum score (rising/stable/declining), a 3-month and 12-month outlook, opportunity windows, and recommended actions. When to use this tool: the user asks what is heating up in a market, wants to time a product roadmap or content calendar, or needs an early read on a sector. Inputs: a sector to monitor and 3-8 keywords defining the watch perimeter. Delivered by Manue, the AI CMO of the Gapup portfolio.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
focusNoOptional context (geography, language target, comparator window, etc.)
sectorYesSector to monitor (e.g. 'B2B SaaS productivity', 'EU fintech', 'climate-tech hardware')
keywordsYes3-8 keywords describing the watch perimeter

Output Schema

ParametersJSON Schema
NameRequiredDescription
kpisNo3-5 headline KPI bubbles
trendsYes5-12 trend cards for the sector
recommendationsNoPrioritised strategic recommendations
executiveSummaryYesBoard-ready sector overview prose
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false. The description adds return structure (5-12 cards, specific fields) and typical use cases. No contradictions. It does not cover rate limits or latency, but the async parameter partially addresses that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the primary purpose. However, the final line about 'Delivered by Manue, the AI CMO' is extraneous and slightly reduces conciseness. Still efficiently communicates key info.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (4 parameters, 2 required, output schema exists), the description adequately covers return format, use cases, and parameters. It is complete enough for an agent to select and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds meaning by explaining sector as 'Sector to monitor' and keywords as '3-8 keywords defining the watch perimeter'. It also mentions inputs in prose, reinforcing schema info. The async and focus parameters are well-described in schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's verb ('monitor') and resource ('emerging trends, regulatory shifts and adoption signals for a given market sector'). It details output (5-12 trend cards with momentum score, outlook, opportunity windows, actions), which distinguishes it from sibling tools like 'competitive_deep_dive' or 'market_entry_strategist'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly specifies when to use this tool: when the user asks what's heating up, wants to time a roadmap/content calendar, or needs an early read on a sector. This provides clear context, though it does not list when not to use or provide explicit alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ugc_moderation_classifierA
Read-onlyIdempotent
Inspect

Multi-language UGC content moderation for marketplaces, social platforms and comment systems. Detects policy violations in text content across 9 policies and 12 languages without external API calls.

Policies checked: • hate — hate speech, slurs, dehumanization (50+ terms × 12 languages) • sexual — explicit sexual content, pornography references, nudity solicitation • violence — threats, weapon references, graphic violence • self_harm — suicidal ideation, self-injury, eating disorder promotion • harassment — doxxing, stalking, cyberbullying, blackmail • scam — phishing, investment fraud, romance scam, lottery fraud • spam — bots, keyword stuffing, excessive caps, emoji storms, suspicious URLs • copyright — piracy, leaked content, serial keys, streaming fraud • minor_safety — grooming signals, CSAM references, minor + adult content combos

Languages: en / fr / de / es / it / pt / nl / zh / ja / ko / ar / ru (auto-detected)

Output includes severity (low/medium/high/severe), confidence (0-100), matched patterns, excerpt, recommended action, age appropriateness (adult/teen/child), and signals.

No API key required. Stateless — no content is stored or logged.

ParametersJSON Schema
NameRequiredDescriptionDefault
langNoLanguage override. If omitted, language is auto-detected.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
contentYesText content to moderate (comment, review, post, chat message).
policiesNoPolicies to check. Default: all 9 policies.
content_typeNoType of content. Affects recommended_action heuristic. Default: comment.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
signalsYes
sourcesYes
violationsYes
lang_detectedYes
quality_scoreYes
age_appropriateYes
content_previewYes
policies_checkedYes
recommended_actionYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds value by specifying statelessness and no logging/storage, and describes output fields like severity and confidence, which align with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points for policies, languages, and output fields. It is front-loaded with the main purpose and concise, with no wasted sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite complexity (9 policies, 12 languages), the description covers all relevant aspects: purpose, policies, languages, output details, and key traits. It is complete for an agent to understand and use the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema covers 100% of parameters with clear descriptions. The description adds context about output fields (severity, confidence, matched patterns, etc.), complementing the schema without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: multi-language UGC content moderation detecting policy violations. It specifies the resource (text content), the action (detects violations), and distinguishes itself from siblings like jailbreak_attempt_detector or bias_amplification_tracker.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context about when to use the tool (any UGC moderation) and notes key traits (no API key, stateless). However, it lacks explicit guidance on when not to use it or alternative tools, leaving the agent to infer from sibling context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upsell_hunterC
Read-only
Inspect

Chasseur d'upsell — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub — Upsell 8 comptes · €127k potentiel · Top 3 : Alan+Qonto+Pennylane · Playbook 5 étapes. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
horizonNo
productYes
accountsYes
targetUpsellEurNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, indicating a safe read operation. The description adds that the tool returns a deliverable and validates inputs server-side, but does not elaborate on latency, cost, or any side effects. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short but includes a verbose reference case (e.g., 'Upsell 8 comptes · €127k potentiel · Top 3 : Alan+Qonto+Pennylane · Playbook 5 étapes') which adds noise without improving clarity. The first sentence could be more front-loaded. Conciseness is acceptable but not exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters, 3 required, deeply nested objects, no output schema), the description is insufficient. It does not explain the output structure, how results are delivered, or what decisions the tools supports. The agent lacks context to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%, and the description does not compensate. It merely states 'send the documented case fields' without explaining the meaning, format, or constraints of required parameters like company, product, or accounts. Parameters with enums and nested objects are left undescribed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns a structured, audited deliverable for upselling, and mentions a reference case. However, it does not explicitly state the core action (e.g., analyzing accounts for upsell opportunities) and fails to distinguish it from sibling tools like cross_sell_reco or renewal_optimizer. The purpose is somewhat clear but lacks precision and differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool vs alternatives. It only mentions that inputs are validated server-side and references a past case, but does not specify context, prerequisites, or exclusions. The agent is left without criteria to choose this over similar tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

usdc_x402_payments_intelA
Read-only
Inspect

Real-time analytics on x402 protocol USDC micropayments for MCP endpoints on Base network. Unique competitive advantage: aggregates internal production telemetry (our own traffic data) with on-chain USDC Transfer events and Bazaar marketplace listings — data no external competitor can access. Four modes: (1) facilitator_stats — Coinbase x402 facilitator settlement statistics (volume, count, top payees/payers). Uses Coinbase CDP API if COINBASE_X402_API_KEY is set; falls back to Base mainnet RPC scan of USDC transfers to known facilitator addresses. (2) endpoint_intel — Per-MCP-endpoint analytics: tx count, USDC volume, unique callers, success rate, catalog size. For gapup-mcp.io endpoints: reads internal JSONL telemetry (richest data source, unique). (3) agent_caller_profile — Anonymous profile of a calling agent wallet: tx count, USDC spent, top endpoints, inferred persona (depth-seeker / bulk-scanner / generalist / researcher / explorer). Wallet anonymised via SHA-256. (4) price_radar — USDC price distribution by tool category (data_lookup / synthesis / compliance / competitive) from Bazaar + internal catalog. Returns median, P25, P75. Network: Base mainnet. USDC contract: 0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913. Cache: 30 min LRU. Timeout per source: 8s. Optional env: COINBASE_X402_API_KEY (higher-fidelity facilitator stats).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesAnalytics mode: facilitator_stats=network-wide settlements | endpoint_intel=per-URL analytics | agent_caller_profile=per-wallet analytics | price_radar=price distribution by category
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
categoryNoTool category for price_radar mode. Defaults to all.
period_daysNoLookback window in days (5-90, default 30)
endpoint_urlNoMCP endpoint URL for endpoint_intel mode (e.g. https://mcp.gapup.io/mcp)
wallet_addressNoEVM wallet address for agent_caller_profile mode (0x...)

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
sourcesYes
price_radarNo
quality_scoreYes
endpoint_intelNo
facilitator_statsNo
agent_caller_profileNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds behavioral context beyond annotations: caching (30 min LRU), timeout (8s per source), optional env var, anonymization via SHA-256, fallback behavior. No contradiction with annotations (readOnlyHint, etc.).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured: opening sentence, competitive advantage, four modes listed with details, then env/cache/timeout. Slightly lengthy but every sentence adds value; front-loaded with main purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given complexity (4 modes, multiple data sources, env-dependent behavior, caching, timeout), the description covers all essential aspects. Output schema exists, so no need to detail return values. Comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions, but the tool description provides additional meaning: e.g., facilitator_stats uses CDP API if key set else RPC scan, category defaults to 'all', period_days range. This enriches understanding beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states specific verb 'analytics' on 'USDC x402 micropayments' for MCP endpoints, lists four distinct modes with clear explanations, and differentiates from siblings via unique data sources (internal telemetry + on-chain + Bazaar).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for when to use each mode (e.g., facilitator_stats for settlement statistics, endpoint_intel for per-endpoint analytics), but does not explicitly say when not to use this tool versus sibling x402 tools (e.g., x402_payment_fraud_detector).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vendor_esg_blacklist_monitorA
Read-onlyIdempotent
Inspect

As a COO, quickly check if a vendor is blacklisted for ESG non-compliance using CDP and GRI data. Input the vendor's legal name or identifier to receive their ESG risk score, blacklist status, and compliance violations. Returns structured data including CDP disclosure score, GRI alignment, and any regulatory flags. Ideal for vendor due diligence, risk assessment, and sustainability reporting. Keywords: ESG, vendor risk, compliance, CDP, GRI, sustainability, blacklist.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNoReporting year (default: current year)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
vendorIdNoOptional identifier (e.g., LEI, DUNS)
vendorNameYesLegal name of the vendor to check

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
vendorIdNo
warningsYes
griAlignedNo
vendorNameYes
violationsNo
blacklistedYes
esgRiskScoreNo
cdpDisclosureScoreNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, idempotentHint) already declare safe read behavior. Description adds context about returned data (CDP disclosure score, GRI alignment, regulatory flags) and mentions quick response, but does not discuss async behavior or rate limits. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise with 4 sentences covering purpose, input, output, and use case. The opening 'As a COO' is slightly unnecessary but does not harm clarity. Ends with relevant keywords.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (4 parameters, existing output schema), the description adequately covers what the tool does, what inputs are needed, and what kind of outputs to expect. Missing details about async parameter behavior or exact output structure, but output schema exists to fill that gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are already documented. Description loosely refers to input (legal name or identifier) but does not add meaningful constraints, formatting, or relationships beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses specific verb 'check if a vendor is blacklisted for ESG non-compliance', clearly identifies the resource (vendor ESG blacklist) and data sources (CDP, GRI). It distinguishes from siblings like vendor_esg_diversity_scanner and vendor_risk_assessor by focusing on blacklist status with specific frameworks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states ideal use cases: vendor due diligence, risk assessment, sustainability reporting. However, it does not explicitly contrast with similar sibling tools or state when not to use it, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vendor_esg_diversity_scannerA
Read-onlyIdempotent
Inspect

For COOs: scans vendor ESG reports to identify suppliers lacking diversity disclosures in GRI or CDP filings. Input a supplier name or identifier to receive a structured assessment of gender, ethnicity, and board diversity metrics. Returns compliance gaps, missing data flags, and source references from CDP open data and GRI standards. Ideal for vendor risk assessment and ESG compliance tracking.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNoReporting year to check (default: current year)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
supplierIdNoCDP or GRI identifier for the supplier (e.g., CDP company ID)
supplierNameYesExact or partial name of the supplier to scan

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
reportLinksNoURLs to relevant ESG reports
supplierNameYes
complianceScoreYesPercentage compliance with diversity disclosure standards
diversityDisclosuresYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the description's disclosure burden is lower. It adds context about the tool being a scanner that returns structured assessments, compliance gaps, missing data flags, and source references from CDP and GRI. However, it does not mention the async behavior (described only in schema), which is a minor omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the action and target audience. Every sentence provides value, with no wasted words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (as indicated by context), the description covers the key outputs (diversity metrics, compliance gaps, source references) and data sources (CDP open data, GRI standards). It does not mention the async parameter's implications, but overall it provides sufficient context for a read-only scanner tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter (year, async, supplierId, supplierName) having a clear description. The description reinforces that the tool expects a supplier name or identifier, but adds no new semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans vendor ESG reports to identify suppliers lacking diversity disclosures in GRI or CDP filings, and returns structured assessments of diversity metrics, compliance gaps, and source references. This specific verb+resource distinguishes it from siblings like supplier_esg_audit (broader ESG audit) and diversity_inclusion_metrics (internal metrics).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description targets COOs and states it is 'Ideal for vendor risk assessment and ESG compliance tracking,' but does not explicitly mention when not to use this tool or compare it to alternatives such as supplier_esg_audit or vendor_risk_assessor. Usage guidance is implied but lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vendor_managementC
Read-only
Inspect

Gestion des fournisseurs — Gapup agent-payable C-suite expertise (COO). Returns a structured, audited deliverable. Reference case: Qonto (12 fournisseurs · €2.4M/an) — €290k économies identifiées · 4 renegociations prioritaires. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
vendorsYes
objectivesYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true. Description adds that it returns a deliverable and provides a reference case, but does not explain the async parameter or output structure. Minimal context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Relatively short, but includes a reference case example that may not be essential. Front-loaded with title and jargon, could be more focused.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given nested objects and no output schema, the description should detail the deliverable format. It only provides a reference case example, missing details on async behavior and output fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (25%), but description does not explain any parameter (e.g., 'company', 'vendors', 'objectives'). Only generic phrase 'send the documented case fields' – no value addition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states it manages vendors and returns a structured deliverable with savings analysis. However, the use of jargon ('Gapup agent-payable C-suite expertise') and lack of explicit differentiation from sibling tools like 'procurement_spend_optim' makes the purpose slightly unclear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description only mentions input validation, not context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vendor_risk_assessorC
Read-only
Inspect

Évaluateur de risque fournisseurs — Gapup agent-payable C-suite expertise (RISK). Returns a structured, audited deliverable. Reference case: Gapup Hub — 15 fournisseurs · €1.8M spend · 3 critiques · Heatmap + plan de remédiation. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
vendorsYes
riskFrameworkNo
assessmentPurposeNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so the description adds minimal behavioral context beyond 'returns a structured, audited deliverable'. No mention of data access, side effects, or internal processing, but no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with a front-loaded purpose and an illustrative reference case. It could be slightly tighter by removing the French branding phrase, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters, nested objects, and no output schema, the description is incomplete. It does not explain the return format or the meaning of risk frameworks and assessment purposes. The reference case hints at output but is insufficient for full understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'async' described). The description does not explain the meaning of 'company', 'vendors', 'riskFramework', or 'assessmentPurpose' beyond their schema names. The reference case provides an example but not parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a vendor risk assessor returning a structured deliverable, with a reference case. However, it does not differentiate from similar sibling tools like supplier_esg_audit or vendor_esg_blacklist_monitor, which also assess vendor risk but with different focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The description only mentions that inputs are validated server-side, but lacks context for appropriate use cases or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vertical_ai_agent_governanceA
Read-onlyIdempotent
Inspect

Generates a comprehensive vertical AI agent workforce integration plan for CHROs, including governance frameworks, human-AI collaboration metrics, and upskilling recommendations. Inputs: industry vertical, workforce size, and current AI adoption level. Outputs: role-specific AI integration roadmaps, skill gap analysis, and performance benchmarks. Uses O*NET skill taxonomies and Gartner AI adoption trends. For best results with large datasets, pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
industryYes
target_rolesNo
workforce_sizeYes
ai_adoption_levelNo
include_benchmarksNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
skill_gap_analysisNo
integration_roadmapNo
collaboration_metricsNo
governance_recommendationsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, openWorldHint, idempotentHint) already indicate safety and idempotence. Description adds specific outputs, data sources (O*NET, Gartner), and async behavior, providing valuable context beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise single paragraph (~60 words) with clear front-loading of purpose. Could benefit from structure like bullet points, but remains efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (mentioned in context signals), the description covers purpose, key inputs, outputs, and a usage hint. Lacks error handling or corner cases, but is reasonably complete for a 6-parameter tool with output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (17%). Description mentions 'industry vertical, workforce size, and current AI adoption level' but omits target_roles, include_benchmarks, and async (though async is noted in usage tip). Partially compensates but not fully for low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states verb 'generates' and resource 'vertical AI agent workforce integration plan' for CHROs, listing included components. Distinguishes from sibling tools by targeting CHROs and workforce planning, but lacks explicit differentiation from other AI governance tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides context for when to use (with industry, workforce size, AI adoption inputs) and an async optimization tip for large datasets. Does not specify when not to use or list alternative sibling tools, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vuln_exploitability_forecastA
Read-onlyIdempotent
Inspect

As a CTO, assess the exploitability risk of CVEs using EPSS scores and cloud asset exposure data. Input a CVE ID (e.g., CVE-2021-44228) to receive exploitability likelihood, affected cloud services, and threat intelligence context. Returns structured risk metrics for prioritization. Sources: CVE NVD, OpenCVE, GitHub Advisories. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
cveIdYes
cloudProviderNo
includeDetailsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
cveIdYes
statusYes
sourcesYes
warningsYes
epssScoreNo
lastUpdatedNo
cloudExposureNo
epssPercentileNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the safety profile is clear. The description adds valuable behavioral context: data sources (CVE NVD, OpenCVE, GitHub Advisories), output type (structured risk metrics), and a warning about potential timeouts with async usage. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise (two sentences plus a list of sources and an async note). It is front-loaded with the purpose. Every sentence adds value: purpose, required input, output description, sources, and key behavioral hint. Could be slightly more structured (e.g., bullet points) but effectively communicates without verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has annotations (readOnlyHint, idempotentHint) and an output schema (not shown but present), the description adequately covers purpose, sources, input hints, and async behavior. It does not explain all parameters or output format in detail, but the output schema fills that gap. The description is fairly complete for a tool with these structured supports.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (25%), with only 'async' having a description in the schema. The description explains 'cveId' (input a CVE ID) and 'async' (pass true to avoid timeout) but does not mention 'cloudProvider' or 'includeDetails'. While it compensates for two key parameters, the missing explanation for the other two leaves gaps. Baseline 3 is appropriate given the partial compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool assesses exploitability risk of CVEs using EPSS scores and cloud exposure data. It identifies a specific verb ('assess'), resource ('CVE exploitability'), and differentiates from sibling tools like cve_security_lookup and vuln_patch_priority_engine by focusing on exploitability forecasting with cloud context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The description implies it's for exploitability risk with cloud data but does not mention when-not to use it or provide comparisons to similar tools like cve_security_lookup or vuln_patch_priority_engine. The note about async is invocation advice, not usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vuln_patch_priority_engineA
Read-onlyIdempotent
Inspect

As a CTO, quickly prioritize unpatched CVEs by combining exploitability scores (EPSS) with cloud asset criticality. Input a list of CVE IDs and your AWS service types (e.g., EC2, RDS) to receive a ranked patching order with risk scores and estimated cloud impact. Uses public NVD, OpenCVE, and AWS pricing data. Ideal for vulnerability management and cloud security posture improvement.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
cveIdsYesList of CVE identifiers to analyze (e.g., ["CVE-2021-44228", "CVE-2023-3824"])
maxResultsNoMaximum number of prioritized CVEs to return (default: 10)
awsServicesNoAWS service types affected by these CVEs (e.g., ["EC2", "RDS", "Lambda"])

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
prioritizedCvesNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly and idempotent. Description adds value by disclosing data sources (NVD, OpenCVE, AWS pricing) and async behavior. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three well-structured sentences. Front-loaded with purpose, followed by output and context. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema (not shown), the description provides sufficient context for inputs and processing. Lacks specifics on scoring methodology but overall adequate for a read-only tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers all 4 parameters with descriptions (100% coverage). Description restates inputs but does not add new semantic details beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it 'prioritizes unpatched CVEs by combining exploitability scores (EPSS) with cloud asset criticality.' It specifies inputs and outputs, and its role is distinct from siblings like cve_security_lookup or vuln_exploitability_forecast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly targets vulnerability management and cloud security posture improvement. It implies when to use (quick prioritization) but does not provide explicit when-not-to-use or compare with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

weather_climate_intelA
Read-only
Inspect

Physical climate intelligence for insurance underwriting, agritech, logistics, energy trading and ESG/climate risk disclosure. Three modes: (1) forecast — 14-day daily weather forecast with temperature, precipitation, wind and humidity; (2) historical — daily records and monthly aggregates for any date range since 1940, with anomaly detection (P90/P95 heat events, extreme precipitation days); (3) climate_risk — long-term physical risk scoring combining CMIP6 ensemble projections (2020-2050), altitude, FEMA flood zones (US) and historical baselines. Risk dimensions: flood, heat (days >35°C/year), drought (SPI), wildfire, sea-level. Overall score 0-100 (100 = severe). Location: city string or lat/lon coordinates. Sources: Open-Meteo (keyless, global, 1940→2050), Open-Elevation, FEMA NFHL (US), NOAA CDO (optional NOAA_API_KEY env var for US+global station data). SLA: ≤25s p95. Cache: 1h forecast / 24h historical / 7d climate_risk.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYes'forecast' (14 days), 'historical' (date range since 1940), 'climate_risk' (long-term physical risk score)
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
date_toNoISO date YYYY-MM-DD — end of date range (required for historical/climate_risk)
metricsNoWeather metrics to include. Default: all metrics.
locationYesGeographic location. Provide either {city, country?} or {lat, lon}.
date_fromNoISO date YYYY-MM-DD — start of date range (required for historical/climate_risk)

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
statusYes
sourcesYes
forecastNo
locationYes
historicalNo
climate_riskNo
quality_scoreYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Disclosure includes SLA (≤25s p95), cache durations, keyless data sources, and risk dimensions. Annotations are consistent (readOnlyHint=true, destructiveHint=false). No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet dense with useful information, well-organized with clear sections for modes, input, sources, and performance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers all essential aspects: modes, parameters, data sources, SLA, cache, and risk dimensions. An output schema exists, so return values are documented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds value by explaining modes, data sources, and risk scoring details beyond schema descriptions. However, parameters like 'async' and 'metrics' are only covered by schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies the tool provides physical climate intelligence for specific industries, explicitly names three modes (forecast, historical, climate_risk) with clear scope, and distinguishes from siblings by its unique focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context for when to use each mode and location inputs, but does not explicitly state when not to use or mention alternatives among the large sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

web_answerAInspect

Get a direct, cited answer to a question, grounded in live web sources. Returns the answer text plus the sources it was built from. Use when you want a conclusion rather than a list of links.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesThe question to answer
includeTextNoInclude the full text of each cited source (default false)
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the output (answer text plus sources) and grounding in live web sources, giving a solid mental model. It doesn't mention latency or async behavior, but those are covered by the schema's async parameter description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with purpose, and contains zero filler. Every word contributes to understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is relatively simple with clear schema docs. The description covers purpose, output, and usage context. Combined with the 100% schema parameter descriptions, an agent has enough to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already well-documented. The description adds minimal semantic value beyond the schema, only implicitly connecting includeText to 'sources.' Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it gets a 'direct, cited answer' grounded in live web sources, and distinguishes itself from 'a list of links' (web search). This clearly separates it from sibling tools like web_search or web_contents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear guidance on when to use: 'when you want a conclusion rather than a list of links.' This implies when not to use it, though it doesn't name specific alternative tools. Still, the context is clear enough for an agent to differentiate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

web_contentsAInspect

Extract the full text of specific URLs, with optional highlights and a summary. Use when you already know which pages you need, rather than searching for them.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYesThe URLs to extract
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
summaryNoAlso return a short summary of each page
highlightsNoAlso return the most relevant excerpts
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the action (extract full text) and optional features (highlights, summary), but does not mention any potential side effects, rate limits, async behavior, or return format. The description adds some value but lacks deeper behavioral context like latency or response structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with purpose and followed by usage guidance. Every sentence is informative, with no redundant fluff. It is concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, and the description does not explain the return format or error behaviors, which is a minor gap. However, given the simplicity of the tool and complete parameter documentation, the description covers the main functional context well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with each parameter described (urls, async, summary, highlights). The description mentions 'optional highlights and a summary' which maps to the summary and highlights parameters, but adds little beyond the schema. The baseline of 3 applies since the schema already documents all parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Extract' and resource 'full text of specific URLs', clearly stating what the tool does. It also distinguishes itself from sibling search tools by noting 'rather than searching for them'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'Use when you already know which pages you need'. It also contrasts with searching, providing a clear when-not scenario and implying the alternative is web_search tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

webhooks_manageAInspect

Manage HTTP webhook callbacks for async tools (T5/T6 batch flagships). Instead of polling every 5s, register a callback URL — Gapup posts the job result to your endpoint the moment it completes. Supported events: job.completed | job.failed | monitoring.alert | quota.threshold. Modes: register (add endpoint), list (view active webhooks), revoke (soft-delete), test (fire a test payload to verify your receiver), history (last 20 fires). Security: every delivery is signed with HMAC-SHA256 on the body — verify the X-Gapup-Signature header against sha256(secret, body).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo(register) HTTPS/HTTP endpoint that will receive POST callbacks. Must return 2xx within 10s.
modeYesregister — add a webhook endpoint. list — view your active webhooks. revoke — soft-delete a webhook by webhook_id. test — fire a test payload to verify the receiver is alive. history — last 20 delivery attempts for a webhook.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
eventsNo(register, optional) Events to subscribe to. Defaults to all events if omitted.
secretNo(register, optional) A secret string used to sign deliveries with HMAC-SHA256. Store it safely — verify X-Gapup-Signature header on your receiver.
webhook_idNo(revoke / test / history) The webhook_id returned from register.
caller_hashNoOptional caller identity override. If omitted, uses the internal session hash.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description provides rich behavioral details beyond annotations: it explains modes (register, list, revoke, test, history), security (HMAC-SHA256 signature verification), and event types. Annotations are readOnlyHint=false and destructiveHint=false, and the description aligns by showing mutable operations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that efficiently covers purpose, usage, modes, events, and security. It is front-loaded with the core purpose. While somewhat dense, it earns its length with no wasted words. Could be slightly improved by breaking into sections, but still effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, 5 modes, security), the description is complete. It covers all modes, events, parameter usage, and security verification. The presence of an output schema (not shown) explains return values, so the description handles the rest adequately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, baseline is 3. The description adds value by explaining each mode's parameter requirements (e.g., 'webhook_id' for revoke/test/history) and security context (secret for signing). This goes beyond the schema's descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool manages HTTP webhook callbacks for async tools, specifying the resource (webhooks) and verbs (register, list, revoke, test, history). It distinguishes from siblings like competitive_deep_dive by focusing on webhooks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells when to use this tool: 'Instead of polling every 5s, register a callback URL'. It also lists supported events and modes. However, it does not explicitly state when not to use it or mention alternatives, leaving a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

web_search_multilangA
Read-only
Inspect

Multi-language, multi-source web search that goes beyond Anglo-centric results. Supports 15 languages (fr/de/es/it/pt/nl/ja/zh/ko/ar/ru/sv/pl/tr/en) with automatic detection. Aggregates results from Mojeek (independent search engine, multilang) and Wikipedia (native multilang API), with DDG and HN as English-language complements. Returns deduplicated results ranked by cross-engine consensus. Use when you need non-English search results, when DDG fails, or for geographically-biased queries. Phase 2 #7 of the geo/lang expansion plan. Note: Brave/Bing/Searx are blocked from DO IPs — configure AICI_RESEARCH_PROXY_URL for residential proxy.

ParametersJSON Schema
NameRequiredDescriptionDefault
langNo2-letter language code. If omitted, auto-detected from query characters and lexical markers.
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
queryYesSearch query in any language
countryNoISO-3166-1 alpha-2 country code for geographic bias (e.g. FR, DE, JP, BR). Optional.
max_resultsNoMaximum number of results to return (default 10).

Output Schema

ParametersJSON Schema
NameRequiredDescription
queryYes
statusYes
resultsYes
sourcesYes
by_engineYes
lang_usedYes
country_usedNo
quality_scoreYes
total_unique_resultsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses key behaviors beyond annotations: aggregation from Mojeek, Wikipedia, DDG, HN; deduplication; async mode with job polling; proxy configuration for blocked IPs. No contradictions with readOnlyHint and destructiveHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused paragraph with no wasted words. It is front-loaded with purpose, followed by key details (languages, sources, usage guidance, notes). Every sentence serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multi-source, multi-language, async, geo-bias, proxy), the description covers all essential aspects. It mentions output characteristics (deduplicated, consensus-ranked) and provides necessary operational context (blocked IPs, proxy setup).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining lang auto-detection, async behavior (returns job_id), country for geographic bias, and max_results default. However, it doesn't add significant meaning beyond the individual schema descriptions for each parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Multi-language, multi-source web search that goes beyond Anglo-centric results.' It specifies the verb 'search', the resource 'web', and distinguishes itself from other search tools by emphasizing multi-language and multi-source aggregation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance: 'Use when you need non-English search results, when DDG fails, or for geographically-biased queries.' It also mentions limitations (blocked from DO IPs) and configuration requirements, helping the agent decide when to use this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

win_loss_decoderC
Read-only
Inspect

Analyse Win/Loss deals — Gapup agent-payable C-suite expertise (CRO). Returns a structured, audited deliverable. Reference case: Gapup Hub — Win/Loss 32 deals Q1 2026 · Win rate 41% → 68% potentiel · Playbook 8 actions CRO. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
dealsYes
companyYes
productYes
topCompetitorsNo
primaryChallengeNo
salesCycleTargetDaysNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and openWorldHint, so the description adds minimal value. It mentions server-side validation and an 'audited deliverable', but lacks details on rate limits, auth needs, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively concise and front-loads the core purpose, though the inclusion of specific percentages and case details adds slight verbosity without critical value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, nested objects, no output schema), the description omits crucial details such as the return format, required vs optional fields, and how to interpret the structured deliverable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (14%), and the description does not elaborate on parameter meanings beyond 'send the documented case fields'. It fails to compensate for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes win/loss deals and returns a structured, audited deliverable, with a specific reference case. However, it does not distinguish itself from similar sibling tools like deal_coach or competitor_intel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions 'send the documented case fields' but provides no guidance on when to use this tool versus alternatives, nor does it specify when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow_orchestratorA
Read-only
Inspect

Meta-tool that CHAINS multiple MCP tools sequentially into a named workflow — delivering a composite output in a single call. 10 predefined workflows: compliance_full_audit (6 steps: KYC+sanctions+AI_gov+privacy+ESRS+CSRD), deal_due_diligence (7 steps: deep_dive+registry+court+patents+KYC+financials+M&A), market_entry_brief (6 steps: country_study+regulations+procurement+tax+AGOA+market_brief), competitor_intelligence_pack (5 steps: deep_dive+intel+patents+earnings+pitch_deck), esg_360 (5 steps: ESG_audit+carbon+CSRD+ESRS+supplier_esg), ip_freedom_to_operate (4 steps: patent_search+async_deep+IP_audit+competitive), climate_property_assessment (3 steps: climate_risk+real_estate+geo), pharma_target_screen (4 steps: trials+adverse_events+patents+meta_analysis), sanctions_360 (5 steps: KYC+Russian_sec+registry+crypto_wallet+court_filings), talent_market_brief (4 steps: salary+trends+adjacent_roles+skills_taxonomy). Returns steps_executed, consolidated P0/P1/P2 signals, overall_status, estimated_cost_usd, and raw outputs per step. Cache: 1h LRU per (workflow, target). Budget: 60s global timeout → partial if exceeded. Use when an agent needs a composite liverable without orchestrating tools manually.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
paramsNoOptional overrides passed to sub-tools. Keys depend on workflow (e.g., country, sector, role, drug, technology, wallet_address, acquirer).
targetYesThe entity to analyze. A company name for most workflows; location for climate_property_assessment; role+country for talent_market_brief.
workflowYesNamed workflow to execute. Each workflow chains 3-7 tools sequentially.
skip_failed_stepsNoDefault true: continue on step failure. Set false to abort on first error.

Output Schema

ParametersJSON Schema
NameRequiredDescription
targetYes
outputsYes
summaryYes
workflowYes
overall_statusYes
steps_executedYes
total_duration_msYes
estimated_cost_usdYes
consolidated_signalsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral context beyond the annotations: details about caching (1h LRU per workflow and target), 60s global timeout with partial results on exceedance, async mode with job_id polling, and a skip_failed_steps option. No contradiction with readOnlyHint=true (the tool orchestrates reads of other tools) and openWorldHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but front-loaded with the core purpose and workflow list. It includes inline examples and behavioral notes. Some redundancy exists (workflow names listed twice), but overall it's efficient for the complexity. Could be slightly more concise without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, 10 workflows, async option, caching, timeout), the description is comprehensive. It covers all essential behavioral aspects, workflow enumeration, and return structure (mentioned but not detailed due to output schema existing). The description leaves no critical gaps for an AI agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While the input schema covers all 5 parameters with descriptions, the tool description adds contextual meaning: it clarifies that the 'target' parameter varies by workflow (e.g., company, location), 'params' are optional overrides, and 'skip_failed_steps' defaults to true. This supplements the schema's definitions, earning above the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly defines the tool as a meta-tool that chains multiple MCP tools into named workflows for composite outputs. It lists 10 predefined workflows with step counts, distinguishing it from sibling tools that are individual analysis tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states 'Use when an agent needs a composite deliverable without orchestrating tools manually.' It provides clear context for when to use, though it does not explicitly list when not to use or alternatives. The sibling tools include many individual tools, but the meta-tool purpose is well-conveyed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

working_capitalC
Read-only
Inspect

Optimiseur du BFR — Gapup agent-payable C-suite expertise (CFO). Returns a structured, audited deliverable. Reference case: Agicap — BFR optimisation · DSO 52→38j · Cash libéré +€2.8M · 3 quick wins immédiats. Inputs are validated server-side — send the documented case fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
companyYes
industryNo
challengesYes
financialsYes
topCustomersNo
topSuppliersNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint. The description adds that inputs are validated server-side and the tool returns a structured deliverable. There is no contradiction, but the description does not provide deeper behavioral insights beyond what annotations imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short (two sentences plus a case) but mixes French and English. It is not overly verbose, but could be more structured and front-loaded with key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (7 parameters, nested objects, no output schema), the description leaves significant gaps. It does not explain the output format or mention the async parameter. More details are needed for an agent to use it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14%, meaning most parameters lack descriptions. The description says 'send the documented case fields' but does not explain individual parameters or their semantics, failing to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it optimizes working capital (BFR) for CFOs and returns a structured deliverable. It includes a reference case, making the purpose clear. However, it does not explicitly differentiate from sibling tools like working_capital_esg_impact_rater, so it loses one point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide guidance on when to use this tool versus alternatives. It mentions 'send the documented case fields' but lacks explicit when-to-use or when-not-to-use criteria relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

working_capital_esg_impact_raterA
Read-onlyIdempotent
Inspect

As a CFO, assess how ESG factors (Environmental, Social, Governance) influence working capital efficiency using IMF SDR and BIS data. Inputs include company sector, geographic exposure, and ESG risk scores. Outputs provide a quantitative impact rating on working capital metrics like days sales outstanding (DSO) and inventory turnover, alongside IMF SDR-aligned liquidity risk indicators.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
regionYesPrimary geographic exposure (e.g., 'EU', 'APAC')
sectorYesIndustry sector (e.g., 'manufacturing', 'energy')
currencyNoReporting currency (ISO 4217 code, e.g., 'USD', 'EUR')
esgRiskScoreYesAggregate ESG risk score (0-100)
workingCapitalRatioNoCurrent working capital ratio (current assets / current liabilities)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
impactRatingNoESG impact on working capital efficiency (-100 to +100)
esgFactorBreakdownNo
liquidityRiskIndicatorNoIMF SDR-aligned liquidity risk score (0-1)
workingCapitalAdjustmentNoProjected adjustment to working capital ratio (%)
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is clear. The description adds context: it uses IMF SDR and BIS data, and produces a quantitative impact rating alongside liquidity risk indicators. No contradictions. This adds moderate value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loaded with the core purpose, and contains no extraneous text. Every sentence contributes meaningful information about inputs, outputs, and data sources. Efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema (not shown but indicated), so the description does not need to detail return values. It explains the outputs (impact rating on working capital metrics, liquidity risk indicators) and data sources. It does not mention the async parameter, but that is a common meta-parameter understood from the schema. Overall, adequate for an agent to use the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%—every parameter has a description in the schema. The tool description only summarizes inputs (sector, geographic exposure, ESG risk scores) without adding new semantic detail beyond what the schema already provides. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'assess how ESG factors influence working capital efficiency' using specific data sources (IMF SDR, BIS). It names the main inputs and outputs, including 'quantitative impact rating on working capital metrics like DSO and inventory turnover'. This distinguishes it from sibling tools like 'working_capital' which likely lacks the ESG focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by a CFO evaluating ESG impact on working capital, but does not provide explicit guidance on when to use this tool versus alternatives (e.g., 'working_capital', 'supplier_esg_audit'). No exclusion criteria or prerequisites are mentioned, which is a gap given the many related sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

working_capital_fx_hedge_optimizerA
Read-onlyIdempotent
Inspect

For CFOs managing multinational working capital, this tool analyzes real-time ECB and FRED foreign exchange rates to recommend optimal hedging strategies. Input base currency, target currencies, and working capital amounts to receive forward contract suggestions, natural hedge opportunities, and cost-benefit analysis of various hedging instruments (forwards, options, swaps). Outputs include hedge ratios, estimated cost savings, and risk reduction metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
baseCurrencyYesISO 4217 code of the company's functional currency (e.g., 'USD', 'EUR')
riskAppetiteNoCompany's risk tolerance for currency fluctuationsbalanced
timeHorizonDaysNoPlanning horizon in days (default: 90)
targetCurrenciesYesISO 4217 codes of currencies to hedge against (e.g., ['EUR', 'GBP', 'JPY'])
workingCapitalAmountsYesWorking capital amounts in each target currency (e.g., { EUR: 5000000, GBP: 3000000 })

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesNo
warningsNo
recommendationsNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint. The description adds value by specifying that the tool uses 'real-time ECB and FRED' data and outputs 'hedge ratios, estimated cost savings, and risk reduction metrics', which are behavioral details beyond the annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loading purpose and target user. Every sentence is essential, with no fluff. It efficiently conveys inputs, process, and outputs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 params, 3 required, nested objects, output schema), the description covers core functionality well. It explains input purpose and output types. Could be more detailed on how strategies are generated, but the output schema likely fills that gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description mentions the key inputs (base currency, target currencies, working capital amounts) and hints at risk appetite and time horizon (default 90), but does not add significant meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool 'analyzes real-time ECB and FRED foreign exchange rates to recommend optimal hedging strategies', with a specific verb ('analyzes... recommend') and resource ('working capital FX hedge'). It distinguishes from siblings like 'fx_rate' and 'treasury_optimizer' by focusing on multinational working capital hedging.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description identifies the target user ('CFOs managing multinational working capital') and required inputs, but does not explicitly state when to use this tool versus alternatives (e.g., when a simple rate lookup suffices via 'fx_rate'). No when-not or exclusion criteria are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

x402_liquidity_monitorA
Read-onlyIdempotent
Inspect

Monitors real-time x402-USDC liquidity depth across 12 decentralized and centralized exchanges, providing slippage alerts and depth analysis for CFO liquidity risk assessment. Inputs include slippage thresholds and exchange selection; outputs liquidity depth, price impact estimates, and warning flags. Essential for optimizing trade execution and managing liquidity exposure. Keywords: liquidity monitoring, slippage analysis, DEX/CEX depth, x402-USDC pair, CFO financial tooling.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
exchangesNoList of exchanges to monitor (defaults to all 12 if empty)
depthLevelsNoLiquidity depth levels to analyze (percentage from mid-price)
slippageThresholdYesMaximum acceptable slippage percentage (0-100)

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
midPriceNoCurrent x402-USDC mid-price
warningsYes
priceImpactNo
liquidityDepthYes
slippageAlertsYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint as true, so the bar is lower. The description adds real-time monitoring behavior and specifics about outputs (liquidity depth, price impact estimates, warning flags), which provides useful context beyond annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a coherent paragraph of four sentences, front-loaded with the main purpose. Each sentence adds value: purpose, inputs/outputs, use case, and keywords. It is efficient but not overly verbose, appropriate for a tool with moderate complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's real-time monitoring across 12 exchanges and existence of an output schema, the description covers inputs, outputs, and use case sufficiently. It doesn't explain error handling or edge cases, but those are likely covered by the output schema and annotations. Overall complete for its complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description reiterates that inputs include 'slippage thresholds and exchange selection' and outlines outputs, but does not add significant semantic meaning beyond what the schema already provides for each parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'monitors' and the resource 'x402-USDC liquidity depth across 12 exchanges', with specific outputs like slippage alerts and depth analysis. It distinguishes itself from sibling tools like 'usdc_x402_payments_intel' and 'x402_payment_flow_analyzer' by focusing on liquidity monitoring and depth analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the tool is 'essential for optimizing trade execution and managing liquidity exposure' and mentions it is for 'CFO liquidity risk assessment'. It implies use cases but does not explicitly state when not to use or provide alternative tools, though the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

x402_payment_flow_analyzerA
Read-onlyIdempotent
Inspect

As a CTO, analyze USDC payment flows involving x402 addresses to assess counterparty risk, trace transaction paths, and evaluate regulatory exposure. Input a wallet address or transaction hash to receive risk scores, flow diagrams, and compliance flags from Chainalysis and TRM Labs public APIs. Ideal for due diligence, fraud detection, and compliance reporting. Pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
depthNoHops to trace in payment flow
txHashNoUSDC transaction hash to trace
addressYesEthereum wallet address to analyze
includeRiskScoreNoInclude counterparty risk scoring

Output Schema

ParametersJSON Schema
NameRequiredDescription
flowIdNoUnique identifier for this payment flow analysis
statusYes
sourcesNo
warningsNo
riskScoreNoCounterparty risk score (0-100)
complianceFlagsNo
exposureSummaryNo
transactionPathNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readonly and idempotent. The description adds that it uses public APIs (Chainalysis, TRM Labs) and supports async mode. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (3 sentences) and front-loaded with purpose. The 'As a CTO' framing is slightly unnecessary but not detrimental.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, the description adequately covers key inputs, outputs, and behavior. Mentions async option, which is important for timeouts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3. The description adds value by highlighting address/txHash inputs and async option, but doesn't explain depth or includeRiskScore beyond schema defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes USDC payment flows for risk assessment, with specific outputs (risk scores, flow diagrams, compliance flags). It distinguishes from sibling tools (e.g., x402_payment_fraud_detector) by focusing on flow analysis and compliance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates use cases (due diligence, fraud detection, compliance reporting) and advises using async:true to avoid timeouts. It does not explicitly list when to avoid, but context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

x402_payment_fraud_detectorA
Read-onlyIdempotent
Inspect

Risk-focused tool that analyzes x402-USDC payment transactions for fraud patterns using on-chain forensics. Takes a transaction hash or wallet address as input and returns risk scores, suspicious indicators, and historical patterns. Designed for risk management teams to quickly assess payment legitimacy. Includes keywords: fraud detection, USDC risk, blockchain forensics, transaction monitoring. pass async:true to avoid timeout.

ParametersJSON Schema
NameRequiredDescriptionDefault
asyncNoIf true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client timeouts.
walletAddressNo
includeHistoryNo
amountThresholdNo
transactionHashYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
sourcesYes
warningsYes
riskScoreYes
isSuspiciousYes
sanctionsMatchNo
fraudIndicatorsNo
transactionHistoryNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, establishing safety. The description adds useful behavioral context by mentioning async usage to avoid timeout, which goes beyond the annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is adequately sized and front-loaded with the purpose, but includes a keyword list that adds noise without value. It could be more concise by removing redundant keywords.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 parameters, low coverage) and presence of an output schema, the description provides a high-level overview but lacks details on input parameters like includeHistory and amountThreshold. It is minimally viable but not fully complete for a nuanced tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 20% schema description coverage, the description should compensate for undocumented parameters. It mentions 'transaction hash or wallet address' but does not explain includeHistory, amountThreshold, or async (beyond a brief note). The description adds little meaning beyond the schema for most parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it analyzes x402-USDC payment transactions for fraud patterns using on-chain forensics, and it takes a transaction hash or wallet address as input. This specific verb+resource combination distinguishes it from sibling tools like fraud_detector or x402_payment_flow_analyzer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions it's designed for risk management teams to assess payment legitimacy and advises using async to avoid timeout, but it does not explicitly state when to use this tool versus alternatives or provide exclusion criteria. Usage is implied but not differentiated from siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Discussions

No comments yet. Be the first to start the discussion!

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    GTM signal intelligence suite for AI agents. Six tools: hiring signals, tech stack detection, company-to-LinkedIn resolution, ICP scoring, job board scanning, and a combined signals aggregator. Built for outbound sales workflows.
    11
    737
    1
    MIT
  • F
    license
    -
    quality
    C
    maintenance
    Browse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.

View all MCP Servers

Try in Browser

Your Connectors

Sign in to create a connector for this server.

Resources