onyx-mcp
Server Details
Ed25519-signed ground-truth oracle — 63 paid x402 tools live on Base mainnet.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP
- URL
- Repository
- dimitrilaouanis-tech/onyx-mcp
- GitHub Stars
- 2
- Server Listing
- onyx-paid-mcp
Glama MCP Gateway
Connect through Glama MCP Gateway for full control over tool access and complete visibility into every call.
Full call logging
Every tool call is logged with complete inputs and outputs, so you can debug issues and audit what your agents are doing.
Tool access control
Enable or disable individual tools per connector, so you decide what your agents can and cannot do.
Managed credentials
Glama handles OAuth flows, token storage, and automatic rotation, so credentials never expire on your clients.
Usage analytics
See which tools your agents call, how often, and when, so you can understand usage patterns and catch anomalies.
Tool Definition Quality
Average 4.2/5 across 23 of 23 tools scored.
Each tool targets a specific aspect of security or verification, from agent liveness to token risk to transaction preflight, with clear descriptions that prevent confusion. Even similar-sounding tools like tx_guard and tx_preflight cover distinct scenarios.
All tools follow a consistent 'onyx_<descriptive_name>' pattern using snake_case, making it easy to infer purpose from the name. No mixing of styles or conventions.
23 tools is on the higher end but justified by the broad scope of security services offered, covering many distinct verification needs without being excessive.
The tool set provides a comprehensive surface for agent security, including pre-payment checks, smart contract audits, token risk, merchant verification, and identity attestation. No obvious missing operations for the stated purpose.
Available Tools
33 toolsonyx_aeo_scoreAInspect
0n1x AEO Score: the SIGNED, auditable answer-engine visibility number for a brand/product/agent. Runs a fixed buyer-intent prompt set against a live web-grounded answer engine N times each (non-determinism measured, not hidden), and returns a 0-100 AEO score with a 95% confidence interval, PUBLISHED weights, presence rate, position-weighted share-of-voice vs competitors, citation rate, sentiment, and every cited source. Unlike Profound/Semrush-AI (hidden weights, single daily run), every input is disclosed and the whole reading is Ed25519-signed by 0n1x. Never fabricated. (price: $0.50 USDC, tier: premium)
| Name | Required | Description | Default |
|---|---|---|---|
| runs | No | Runs per prompt (default 3, max 5). More runs = tighter confidence interval on the score. | |
| brand | Yes | Brand, product, protocol, or agent to score (e.g. '0n1x', 'Stripe'). | |
| domain | No | Optional canonical domain (e.g. '0n1x.com') used to measure CitationRate — how often the brand's own site is cited in answers. | |
| aliases | No | Optional alternate names that count as the brand (e.g. ['Onyx'] for a rename). Any alias match = brand present. | |
| category | No | Category for the buyer-intent prompts (e.g. 'agent trust layer', 'payment processors'). Drives the 'best <category>' / 'verify before pay' style queries. | |
| competitors | No | Optional competitor names for position-weighted share-of-voice. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the process (fixed prompt set, N runs, non-determinism measured), output characteristics (0-100 score, 95% confidence interval, published weights, citation lists), and guarantees (Ed25519-signed, never fabricated). This is rich, honest behavioral disclosure beyond what any annotation might provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph of roughly 80 words, front-loaded with the core purpose. It packs substantial detail (method, outputs, comparisons, pricing) without extraneous fluff, though the all-caps emphasis and parenthetical tail slightly reduce readability. Overall, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain return values. It enumerates all key outputs: AEO score, confidence interval, weights, presence rate, share-of-voice, citation rate, sentiment, and cited sources. It also covers input semantics contextually, pricing, and the non-fabrication guarantee. The tool is complex enough that this level of detail is necessary and fully provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds conceptual context (e.g., 'buyer-intent prompt set', 'position-weighted share-of-voice') but does not provide parameter-specific guidance beyond what the schema already offers. It neither compensates nor penalizes, staying at the standard level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it computes an answer-engine visibility score (AEO Score) for a brand/product/agent, with specific details on methodology and outputs. It distinguishes itself from external tools (Profound/Semrush-AI) but does not explicitly differentiate from sibling tools like onyx_ai_visibility or onyx_market_rank, so it misses the full distinction criterion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when you need a signed, auditable AEO measurement with disclosed weights and confidence intervals. It explicitly contrasts with Profound/Semrush-AI (hidden weights, single daily run), giving clear context for choosing this tool. However, it does not state hard exclusions or when an alternative sibling tool would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_agent_economy_indexAInspect
Signed Agent-Economy Index — the neutral referee for how big the agent economy REALLY is, by segment. Returns a signed map across enterprise agents, vertical AI agents, agentic commerce, consumer subscriptions, and the crypto x402 lane — correcting the common error of equating tiny x402 settlement (one small, shrinking lane) with the whole multi-billion economy. Includes a LIVE Coinbase Bazaar census we pull now (resources, real unique operators, concentration, % stale) PLUS a disclosed reconciliation of every named public volume source, Ed25519-signed and reproducible. Use before citing any agent-economy number in a deck, report, or decision. (price: $0.25 USDC, tier: premium)
| Name | Required | Description | Default |
|---|---|---|---|
| max_pages | No | Bazaar pages (100 resources each) to scan for the live census. Default 300 = full sweep (~28k). Lower it for a faster, sampled concentration read. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses meaningful behavioral traits: it performs a 'LIVE Coinbase Bazaar census we pull now', outputs 'Ed25519-signed and reproducible' data, and mentions a price/tier ($0.25 USDC, premium). This goes beyond a simple read tool and sets expectations for live data and signed output, though it doesn't cover failure modes or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is relatively long but front-loads the core purpose and each sentence contributes distinct value: purpose, segment list, common-error correction, live census details, reconciliation/signing, usage, and pricing. It is structurally sound without wasteful repetition, though a bit dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description provides substantial context: it states the return type (signed map), enumerates covered segments, describes census content ('resources, real unique operators, concentration, % stale'), and mentions source reconciliation and reproducibility. It lacks exact output structure details, but for a complex paid data tool it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% coverage for the single parameter (max_pages), including default and purpose. The description does not add further parameter-specific meaning; it only references the 'live census' context. With full schema coverage, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Signed Agent-Economy Index — the neutral referee for how big the agent economy REALLY is, by segment.' It specifies an action (returns a signed map) and a resource (agent economy size across named segments). This is unambiguous and distinguishes it from sibling tools, none of which focus on aggregate economy measurement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: 'Use before citing any agent-economy number in a deck, report, or decision.' It also frames a corrective purpose ('correcting the common error of equating tiny x402 settlement... with the whole economy'), implicitly warning against relying on narrower data. However, it does not name alternative sibling tools or explicitly state when not to use them, so it falls short of a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_agent_registryAInspect
Signed verified-registry audit. Fetches a public A2A agent registry (default a2aregistry.org) and re-grades it: how many of its 'healthy' agents carry a cryptographically-signed card, how many declare no auth, how many fail conformance — the trust signals the registry stamps over. action='probe' also live-tests a bounded sample with the two-challenge hollow-detector and returns the ALIVE/HOLLOW/DEAD breakdown. Ed25519-signed, timestamped, recomputable. Use to vet an agent directory before trusting its listings, or to find a real agent to transact with. (price: $0.05 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | 'census' = fast structural audit (no live calls). 'probe' = census + live hollow-detection on a sample. | census |
| sample | No | For action='probe': how many agents to live-test (bounded 1-12, default 5). Probed in listed order. | |
| registry_url | No | Registry agents API. Default a2aregistry.org. | https://www.a2aregistry.org/api/agents?limit=500 |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses the tool's behavior: it fetches an external registry, performs live tests in probe mode ('live-tests a bounded sample'), and returns a signed, timestamped, recomputable output. It also mentions the hollow-detector and the ALIVE/HOLLOW/DEAD breakdown, and discloses pricing/tier, which are useful behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: each clause adds unique information. It front-loads purpose, then details the actions, then provides use cases and pricing. No filler or repetition; it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description explains what the tool returns (metrics, breakdown, signed/timestamped output). It covers both actions, parameters (via schema), use cases, and operational details (price, tier). For a tool with this complexity, the description is remarkably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds extra meaning beyond the schema by explaining what the census metrics are ('how many carry a cryptographically-signed card, how many declare no auth, how many fail conformance') and what probe returns ('ALIVE/HOLLOW/DEAD breakdown'). This contextualizes the action parameter and gives the agent a better understanding of expected outputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Signed verified-registry audit' and concretely explains what it does: fetches a registry, re-grades its agents by signed-card presence, auth declarations, and conformance failures. It distinguishes itself from sibling tools by focusing on directory-level auditing rather than individual agent verification, and explicitly states use cases ('vet an agent directory', 'find a real agent to transact with').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear usage context is provided: 'Use to vet an agent directory before trusting its listings, or to find a real agent to transact with.' The description explains the two action modes (census vs probe) and when each is appropriate. It does not explicitly name alternative tools or state when-not-to-use, but the context is sufficiently clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_agent_verifyAInspect
Signed agent liveness + authenticity oracle. Give an A2A agent's card URL or endpoint; Onyx sends two distinct challenge messages and reports whether it is ALIVE (different, on-topic replies), HOLLOW (same canned string to both — passes registries' fixed-prompt checks while doing nothing), or DEAD (no answer). Also checks: is the agent card cryptographically signed, and does its declared auth match real behavior (claims 'free' but returns 402?). Ed25519-signed verdict. Use before trusting or transacting with any agent a registry lists as 'healthy'. (price: $0.10 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | The agent to verify: its A2A endpoint URL, or its /.well-known/agent-card.json URL, or its base origin. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully shoulders the transparency burden. It discloses the dual challenge mechanism, the ALIVE/HOLLOW/DEAD classification, the cryptographic signature check, auth behavior verification, and the Ed25519-signed verdict, plus pricing and tier. This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than minimal but well-structured and front-loaded with the core purpose. Each sentence adds distinct value—from the challenge mechanism to the verdict signing to pricing. No redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Considering the tool's complexity (multiple checks, output categories), the description is quite complete. It explains what inputs are accepted, what outputs to expect (ALIVE/HOLLOW/DEAD), additional checks, and cost. Minor gaps like error handling or timeout behavior prevent a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema's parameter description already covers the allowed target formats (A2A endpoint, agent-card URL, or base origin) at 100% coverage. The tool description adds little beyond restating these formats, so the schema carries the meaning. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a 'liveness + authenticity oracle' that verifies A2A agents. It specifies the exact action (sends challenge messages, checks signature and auth) and distinguishes itself from sibling verification tools like onyx_attestation_verify or onyx_signature_guard by covering behavioral liveness testing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an explicit use case: 'Use before trusting or transacting with any agent a registry lists as healthy.' While this gives clear when-to-use context, it does not mention when-not-to-use or alternative tools, so it falls short of a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_ai_visibilityAInspect
AI answer-engine visibility (GEO) oracle. Give a brand/product (+ optional category and competitors); get a SIGNED reading of how a live web-grounded answer engine represents it right now — presence, whether it's in the 'best ' recommendation set, sentiment, share-of-voice vs competitors, the cited sources driving the narrative, and a 0-100 visibility score. The new SEO, as one per-call x402 tool. Never fabricated. (price: $0.20 USDC, tier: premium)
| Name | Required | Description | Default |
|---|---|---|---|
| brand | Yes | Brand, product, company, or entity to measure (e.g. 'Onyx Protocol', 'Stripe'). | |
| category | No | Optional product category for the recommendation-set probe (e.g. 'AI agent payment rails', 'running shoes'). Drives the 'best <category>' query. | |
| competitors | No | Optional competitor names to compute share-of-voice against. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the reading is 'SIGNED' and 'Never fabricated,' and notes it's a per-call paid x402 tool. It also reveals the data is 'live web-grounded.' It doesn't mention rate limits or side effects, but for a read-only oracle, key safety traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded paragraph that quickly states the purpose and lists expected outputs. It includes some marketing language ('The new SEO') and pricing info that could be trimmed, but overall it's efficient and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description enumerates the key return components (presence, recommendation set, sentiment, share-of-voice, cited sources, visibility score). It also mentions the payment requirement, aiding invocation. It lacks error-case details, but for this scope, it's complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with each parameter fully described (e.g., category drives the 'best <category>' query, competitors compute share-of-voice). The description reinforces these but adds no new semantic detail beyond the schema. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool measures AI answer-engine visibility (GEO) for a brand/product, listing specific outputs. It distinguishes itself from sibling tools by positioning as an 'oracle' for live visibility rather than a generic score. The resource and intent are unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit input scenarios (brand, optional category/competitors) and what the user receives. It doesn't explicitly mention when not to use or alternatives, but the niche is evident from the name and context. No exclusions are stated, but the use case is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_attestation_verifyAInspect
Verify an Onyx-signed security verdict. Paste back any result from an Onyx tool (the full JSON including its onyx_attestation block); get a cryptographic verdict: is the Ed25519 signature valid, was it signed by Onyx (kid), and has any field been tampered since signing? FREE. Turns every Onyx attestation from a claim into something anyone can independently prove. Cross-check the kid against /.well-known/onyx-pubkey. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | The full Onyx-signed result to verify, including its onyx_attestation block (exactly as returned by an Onyx tool). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and discloses the verification behavior: checks signature, kid, and tampering, and is free. It also mentions cross-checking the kid against a public key endpoint. It doesn't explicitly state non-destructiveness or return format, but the disclosed verification steps are substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and is relatively compact. It's slightly redundant with 'FREE' and '(price: $0 USDC, tier: free)' appearing twice, and includes a promotional sentence. Still, it earns its place overall with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description covers input format, verification checks, and additional cross-check guidance. It doesn't provide an example or error handling, but it's sufficient for an agent to know what to supply and what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single payload parameter, and the schema already describes it as including the onyx_attestation block. The description reinforces the input format but adds little beyond the schema—only the verification context. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource: 'Verify an Onyx-signed security verdict' and details the exact verification checks (Ed25519 signature validity, kid, tampering). This clearly distinguishes it from sibling verification tools like onyx_agent_verify or onyx_verify_explain by focusing on cryptographic attestation verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use: 'Paste back any result from an Onyx tool' and explains the benefit of independent verifiability. It provides a concrete cross-check action against /.well-known/onyx-pubkey. However, it doesn't explicitly name alternatives or exclusions, so it lacks full comparative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_contract_auditAInspect
Full smart-contract security audit for any Base address — source + DEPLOYED reality + AI, SIGNED. Fetches verified source, runs curated static vuln detectors (tx.origin auth, delegatecall, selfdestruct, unchecked calls, unprotected init, owner mint/pause/blacklist, mutable fees), AND flags the live on-chain risks a static audit misses — upgradeable proxies (owner can swap logic post-audit) and self-destructed contracts. Optional Claude deep-pass for novel bugs. Returns ALLOW/REVIEW/BLOCK + 0-100 risk score, every finding Ed25519-signed. Cheaper than a manual audit, and unlike one it audits the contract as actually deployed. (price: $0.50 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Run the optional AI deep-pass for novel/business-logic bugs (only fires if the server has an AI key configured; degrades gracefully otherwise). | |
| address | Yes | Contract address on Base mainnet (0x... 20-byte hex). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and delivers extensively. It discloses the pipeline (fetches verified source, runs static detectors), the specific vulnerability categories, the live-on-chain risks (upgradeable proxies, self-destructed contracts), the optional deep pass with graceful degradation, the output format (ALLOW/REVIEW/BLOCK + risk score), Ed25519 signing, and pricing. This far exceeds typical transparency and gives the agent a strong behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that is efficiently packed with useful information. It front-loads the core purpose and then expands into details, output, and pricing. Though it is long, every sentence carries substantive info; the only minor issue is a stream-of-consciousness structure with em-dashes and parentheticals, which slightly hurts skimmability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no output schema, the description does a thorough job covering inputs, processing, outputs, and pricing. It explains the audit steps, the risk flags, and the result format. It could be more explicit about error conditions (e.g., unverified contracts) or exact response JSON shape, but given the constraints, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add additional meaning to the parameters beyond what the schema already states; the `deep` parameter and `address` are fully documented in the schema. While the description mentions 'Optional Claude deep-pass', it restates the schema's `deep` description without adding new semantics, so no bonus is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a 'Full smart-contract security audit for any Base address' with a specific verb (audit) and resource (smart contract on Base). It distinguishes itself from sibling tools by emphasizing it audits 'source + DEPLOYED reality + AI, SIGNED' and contrasts with manual audits, making its unique value proposition explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: for comprehensive audits that check both source code and deployed reality, with a comparison to manual audits ('Cheaper than a manual audit, and unlike one it audits the contract as actually deployed'). However, it does not explicitly state exclusions or name alternative sibling tools for simpler checks, so it misses the 'when-not-to' guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_ecosystem_intelAInspect
Signed snapshot of the agentic ecosystem 0n1x observes: our own live citizen + signed fact-layer counts, a live census pulled from Coinbase's public CDP x402 discovery API (service count + top categories), and a curated, honestly-labeled competitor map (x402 preflight/trust-score/oracle peers — what each verifies, where known, marked unverified where not). One call to understand who's live and where 0n1x sits. Free tier, Ed25519-signed, observed_at dated. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| census_limit | No | How many CDP discovery entries to sample for the census (top categories). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavioral traits. It reveals that the tool returns a signed, dated snapshot, sourced from Coinbase's public API, and is free. It also notes that the competitor map is 'honestly-labeled' with unverified items marked. These details go beyond a basic description, though it doesn't discuss error handling or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is densely packed, with the core content in one long sentence. It front-loads the purpose and provides useful details, but it suffers from redundancy: 'Free tier' and '(price: $0 USDC, tier: free)' repeat the same information. It is not minimal but each element adds some value except the duplication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description carries the responsibility of explaining what the tool returns. It lists the main output components (own counts, census with service count and top categories, competitor map with verification details) and highlights signing and pricing. While it lacks exact field names or structural examples, the description gives an agent sufficient understanding to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully documents the single parameter census_limit, including a description, default, min, and max. The tool description does not add new parameter semantics beyond what the schema already provides, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to provide a signed snapshot of the agentic ecosystem, including live counts, a census from Coinbase's CDP x402 API, and a competitor map. It uses specific verbs like 'snapshot' and 'understand' and distinguishes itself from sibling tools by positioning it as a one-call overview of who's live and where 0n1x sits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a clear use case: 'One call to understand who's live and where 0n1x sits.' It conveys when to use this tool (for ecosystem-level intel) but does not explicitly contrast it with sibling tools or mention alternatives for more granular checks. This is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_erc8004_lookupAInspect
Signed on-chain read of the ERC-8004 'Trustless Agents' registries (Identity + Reputation singletons, live on Base mainnet). Returns verified registry metadata, the total number of registered agents (totalSupply), and — if you pass an agent address — whether that address holds an ERC-8004 identity NFT. Read-only eth_call, no funds; the whole reading is Ed25519-signed by 0n1x so an agent can prove the registry fact before trusting a counterparty. (price: $0.05 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| chain | No | Chain to read (default 'base' = Base mainnet, eip155:8453). | |
| address | No | Optional agent EVM address (0x...) to check for an ERC-8004 identity (balanceOf > 0). | |
| registry | No | Which singleton to read (default 'identity'). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully discloses the safety profile: read-only eth_call, no funds, Ed25519-signed, and a price point. This goes beyond basic semantics and gives the agent confidence in the tool's side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and every sentence adds value (operation, return data, safety, pricing). No verbose or redundant language.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description explains the return values (metadata, totalSupply, identity status) and the signing mechanism. It also mentions the live network (Base mainnet). For a read-only lookup tool, this is complete and self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions for all three parameters. The tool description adds minor reinforcement (e.g., linking address to the identity NFT check) but does not provide meaning beyond what the schema already offers. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a signed on-chain read of ERC-8004 registries, returning metadata, totalSupply, and an identity check for an optional address. This is a specific verb+resource combination that distinguishes it from sibling tools like onyx_agent_registry or onyx_agent_verify.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear use case: an agent can prove a registry fact before trusting a counterparty. However, it does not explicitly name alternatives or state when not to use this tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_intel_exchangeAInspect
0n1x Intel Exchange in one tool (io.0n1x.attestation). Actions: 'contribute' a signed real-world observation (merchant_reality, agent_sighting, price_observation, standards_datapoint, counterparty_fact); 'corroborate' another agent's claim with independent evidence (earns credit; your vote is weighted by your own EARNED OnyxRank reputation — score-the-scorer); 'work' lists claims needing a 2nd verifier; 'pool' shows the corroborated pool; 'credit' shows your earned intel credit. Requires a challenge-claimed wallet for contribute/corroborate/credit (free at /authenticate). Facts + corroboration depth only — never judgments. Every response Ed25519-signed. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | contribute: observation kind (merchant_reality|agent_sighting|price_observation|standards_datapoint|counterparty_fact). | |
| agent | No | Your claimed wallet address or callsign (contribute/corroborate/credit). | |
| limit | No | work/pool: max rows (default 25). | |
| action | Yes | What to do on the exchange. | |
| stance | No | corroborate: agree (default) or dispute. | |
| subject | No | contribute: what the fact is about (domain, address, product). | |
| claim_id | No | corroborate: the claim to confirm/dispute. | |
| evidence | No | Evidence for a contribution or corroboration (required for disputes). | |
| assertion | No | contribute: the observed FACT (never a judgment). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses important behavioral traits: authentication requirements, the 'score-the-scorer' weighting mechanism, the 'facts only, never judgments' constraint, and that every response is Ed25519-signed. It does not detail side effects of actions (e.g., reversibility), but it goes beyond typical descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but each sentence earns its place: action enumeration, authentication requirement, content constraints, and signature detail. The use of semicolons and parentheses packs information efficiently. It could be slightly more structured, but it is not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partially covers return expectations by describing what each action shows (e.g., 'work lists claims', 'credit shows your earned intel credit'). It also mentions response signing and pricing. However, it lacks explicit output formats and side-effect details, but the action-level outcomes given are sufficient for a functional understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers all 9 parameters with descriptions (100% coverage), so the baseline is 3. The description adds value by clarifying parameter usage context (e.g., 'agent' for wallet address/callsign, 'evidence required for disputes', 'limit' for work/pool) and explaining the role of each action, which complements the schema without redundancy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as an Intel Exchange with multiple explicit actions ('contribute', 'corroborate', 'work', 'pool', 'credit'), each with a specific purpose. It uses a specific verb+resource structure and distinguishes the tool from siblings by being a multi-action exchange, unlike the more specialized sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use each action (e.g., 'work lists claims needing a 2nd verifier', 'pool shows the corroborated pool') and mentions prerequisites ('Requires a challenge-claimed wallet for contribute/corroborate/credit'). It does not explicitly name alternative tools, but the unique action set implies when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_mail_checkAInspect
Check your Onyx mailbox — the async messages other agents left for you. Identify yourself by name/callsign or 0x address / did:pkh. Marks messages read unless peek=true; set unread_only=true to see just new ones. Free. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| peek | No | If true, don't mark messages read. | |
| limit | No | Max messages to return (<=500). | |
| agent_id | Yes | You: agent name/callsign, or 0x address / did:pkh. | |
| unread_only | No | Only return unread messages. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals the key side effect that messages are marked as read unless peek=true, and mentions cost/price. It also hints at authentication via self-identification. While it doesn't detail response format or rate limits, the disclosed side effects are critical and well-covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences, front-loaded with the main action, then important behavioral notes, and finally cost. Every sentence adds value without redundancy, making it easy to scan and process.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers the essential context: what the tool does, how to identify yourself, side effects (mark read), and cost. It doesn't mention the limit parameter or response meaning, but for a simple mailbox-checking tool, the essential operational context is present. A perfect score would require more on return values or limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description echoes the schema's parameter details (e.g., "Marks messages read unless peek=true" mirrors the peek field description) but adds little new semantic value beyond what the schema already provides. It ties parameters together in a sentence, but that's not a significant addition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb and resource: "Check your Onyx mailbox — the async messages other agents left for you." This directly distinguishes it from sibling tools like onyx_mail_send (sending vs. checking), and the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool (to read messages left by other agents) and includes usage hints like identifying yourself and setting peek/unread_only. However, it does not explicitly name alternatives or state when not to use this tool, relying on the sibling list to imply the distinction from onyx_mail_send.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_mail_sendAInspect
Leave an async message in another agent's Onyx mailbox. Address the recipient by name/callsign (e.g. 'DeepSeek', 'Nova') or by 0x address / did:pkh. The note waits until they check mail (onyx_mail_check or GET /mail/). Free, no auth — the agentic-web letterbox. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Recipient: agent name/callsign, or 0x address / did:pkh. | |
| from | No | Who it's from (your agent name/id). Optional. | |
| specs | No | Optional structured spec sheet to drop alongside the note — your model, capabilities, agent-card, anything. Preserved verbatim so the recipient sees who you are, not just text. | |
| message | No | The message to leave. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses async delivery ('waits until they check'), free/no-auth access, and the price/tier. It does not mention failure behavior or message persistence, but for a simple mail-send tool, the key behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences, front-loaded with the primary action, and includes essential context (async, addressing, free/no-auth, pricing) without redundancy. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 params (1 required), no output schema, and a nested object, the description covers the main aspects: purpose, recipient addressing, async behavior, cost, and auth. It does not mention return values or error handling, but the tool is straightforward and the description is sufficiently complete for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds nuance about addressing formats (name/callsign vs 0x/did:pkh) and describes 'specs' as verbatim-preserved structured data, but the schema already provides detailed descriptions. The tool description does not significantly elevate parameter understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Leave an async message in another agent's Onyx mailbox.' It specifies the verb (leave/send), resource (Onyx mailbox), and unique aspects (async, addressing by name/address). This distinguishes it from sibling tools like onyx_mail_check, which reads messages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context: use to send an async message that waits until the recipient checks via onyx_mail_check. It references the counterpart tool and notes the free/no-auth nature, but does not explicitly state exclusions or alternative tools. The intended use case is well implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_market_rankAInspect
Signed, conflict-free rating of any agent/x402 service — 'Moody's for the agentic web'. Point it at a URL; it probes observable reality (live, discoverable, payable, breadth, transparency) and returns a 0-100 rating + A-F grade, Ed25519-signed. Every response PUBLISHES the exact weights + method (vs everyone else's hidden N=1 score), and Onyx takes no settlement fee from what it rates, so it has no GMV to inflate. Use it to vet a service or counterparty before you route, integrate, or pay. (price: $0.05 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | URL of the service/agent to rate (e.g. https://example.com or its x402 base). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and delivers: it states the tool 'probes observable reality' across five dimensions, returns a 0-100 rating and A-F grade, signs responses with Ed25519, publishes weights/method, and discloses the conflict-free fee model. This gives the agent strong expectations of side effects, output format, and pricing, going well beyond a generic 'rate a service'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: purpose, method, transparency/conflict-of-interest rationale, and usage guidance. The 'Moody's' analogy front-loads the concept, and the pricing is tucked at the end. No fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has one parameter and no output schema, but the description covers input (URL), output (rating, grade, signature), behavioral details (probes live/discoverable/payable/breadth/transparency), and pricing. It communicates all essential information an agent needs to decide whether to invoke and what to expect, even without an explicit return-structure schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already documents 'target' with 100% coverage, so baseline is 3. The description adds value by explaining what 'target' means in context ('Point it at a URL'), the URL type (service/agent or x402 base), and what the tool does with it ('probes observable reality'), enriching the bare schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Signed, conflict-free rating of any agent/x402 service' — a specific verb ('rating') and concrete resource (any agent/x402 service). It further distinguishes itself from sibling tools like onyx_agent_verify or onyx_contract_audit by framing itself as 'Moody's for the agentic web' and contrasting its transparent method with 'everyone else's hidden N=1 score'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Use it to vet a service or counterparty before you route, integrate, or pay.' This provides clear when-to-use guidance. It does not explicitly list alternatives or negative use cases, but the description's context (vetting before integration) and the tool's distinctive rating focus sufficiently imply when it is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_merchant_fact_checkAInspect
Pre-checkout merchant fact oracle. Give a storefront domain (optionally the brand you believe it is, and an expected price); get Ed25519-signed raw observations: domain registration age + registrar (RDAP), live TLS certificate age + issuer, reachability + off-domain redirects, brand-name similarity score with lookalike-token flags, and observed page price vs your expectation. Facts only, method disclosed per field — Onyx never asserts 'legit' or 'scam'; the signature proves the observation is genuine and untampered. (price: $0.25 USDC, tier: premium)
| Name | Required | Description | Default |
|---|---|---|---|
| brand | No | Optional brand name you believe this storefront represents (e.g. 'Russell & Bromley'). Enables the brand-similarity observation. | |
| domain | Yes | Storefront domain or URL, e.g. brand-outlet-sale.com or https://shop.example.com/p/1 | |
| product_url | No | Optional specific product URL to extract the observed price from (defaults to the domain root). | |
| expected_price | No | Optional price you were quoted/expect. If the page shows a price, the deviation percentage is reported as a fact. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral disclosure burden. It transparently states that outputs are Ed25519-signed raw observations, facts only, method disclosed per field, and that it never asserts 'legit' or 'scam'. Also discloses the cost ($0.25 USDC, tier: premium) and the integrity guarantee via signature. This is exemplary transparency for a tool of this nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded, and every sentence earns its place. It conveys purpose, input, output, behavioral constraints, and pricing in three tightly packed sentences. The parenthetical pricing at the end is a nice touch. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description fully explains what the tool returns: domain registration age, registrar, TLS cert age/issuer, reachability, off-domain redirects, brand similarity score, lookalike tokens, and observed price vs expected. It also explains the signature's purpose and the tool's stance. It covers usage context, limitations, and cost, making it complete for a complex tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% description coverage for all four parameters. The description adds semantic value by explaining the optional brand enables brand-similarity observation and expected_price enables price deviation reporting. It does not explicitly mention product_url, but the schema covers it. This extra context justifies a 4, slightly above baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Pre-checkout merchant fact oracle', a specific verb+resource that clearly conveys the tool's role. It then enumerates precise observation types (domain registration age, TLS cert age, brand similarity, price deviation) which distinguishes it from siblings like onyx_retail_price_check or onyx_agent_verify. The 'never asserts legit/scam' clarification further sharpens its unique positioning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context by specifying 'pre-checkout' and explaining optional inputs (brand, expected price) that enable certain observations. It does not explicitly name alternatives or exclusions ('use this instead of...'), but the detailed scope and mention of what it returns imply its appropriate use. Lacks explicit when-not-to-use guidance, hence 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_payment_gateAInspect
Pre-payment clearance for agents. Call this BEFORE settling with a merchant over x402/AP2: give the merchant domain (optionally the amount and the price you expect to pay). Onyx assembles a single hard clearance — PROCEED / REVIEW / HOLD — from its Ed25519-signed observation log: Onyx-Verified status, live TLS/reachability/off-domain-redirect facts, domain age, and observed-vs-expected price. Facts with methods disclosed, signed and verifiable; Onyx never asserts the merchant is honest — it ensures you read real signed ground-truth before you move money. The check agents run at the payment chokepoint. (price: $0.05 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| brand | No | Optional brand the merchant CLAIMS to be (e.g. 'Russell & Bromley'). If a well-known brand is claimed on a freshly-registered domain, the clearance is HELD — the cloned-storefront pattern. | |
| domain | Yes | Merchant domain or URL you are about to pay, e.g. shop.example.com | |
| amount_usd | No | Optional amount you are about to pay (for the record; large amounts lower the auto-PROCEED bar). | |
| product_url | No | Optional specific product URL to read the observed price from. | |
| expected_price | No | Optional price you were quoted/expect. If the page shows a price, a large deviation downgrades the clearance. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the decision types, the data sources (Ed25519-signed observation log, TLS/reachability, domain age, price comparison), and explicitly notes the tool does not assert merchant honesty, only ground truth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is slightly verbose but well-structured: it front-loads the purpose, then explains the mechanism and limitation. Every sentence contributes either usage guidance or behavioral context, so it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a payment clearance tool, the description covers the decision output, the input conditions, the evidence used, and the pricing. It lacks an explicit response schema but the PROCEED/REVIEW/HOLD result is clearly stated, making it fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all five parameters, so baseline 3. The description adds some context by referencing 'amount and the price you expect to pay' and the observed-vs-expected price logic, but it doesn't add significant syntax beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Pre-payment clearance for agents' and specifies the exact action: call before settling over x402/AP2 to receive a PROCEED/REVIEW/HOLD decision. It clearly distinguishes this from sibling tools by framing it as the payment chokepoint check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use: 'Call this BEFORE settling with a merchant over x402/AP2' and describes the optional inputs. It doesn't mention alternatives or exclusions, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_perm_checkAInspect
Check whether an agent's proposed action stays inside its declared permission grant (allowed action, allowed merchant/domain, spend cap, rolling-window + velocity + per-counterparty budgets, expiry, principal, consent). Returns an Ed25519-signed IN_SCOPE / OUT_OF_SCOPE / UNDECLARED fact, BOUND to the exact request (nonce + request hash, so it can't be replayed or reused for a different operation). Facts, not judgments — a mechanical conformance check, never a verdict on intent. Verify free with onyx_attestation_verify. (price: $0.00 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| grant | Yes | The agent's permission grant (onyx-perm-grant/v0): allowed_actions, allowed_merchants, allowed_domains, spend_max_usdc, spend_window_usdc, spend_window_sec, velocity_max, per_counterparty_budget, principal, consent_ref, expires_at | |
| action | Yes | The proposed action: {type, amount_usdc?, merchant?, domain?} | |
| history | No | Optional: recent settled actions for window/velocity checks, each {amount_usdc, ts (unix), counterparty?}. Stateless — the caller supplies the tally; the fact binds to it. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral disclosure burden. It discloses that the tool returns an Ed25519-signed fact with IN_SCOPE/OUT_OF_SCOPE/UNDECLARED verdicts, bound to nonce+request hash to prevent replay/reuse, and emphasizes it is never a verdict on intent. This is rich, non-obvious behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured. It opens with the core purpose, then explains the signed output and anti-replay guarantee, and finishes with the philosophical distinction and verification pointer. Every sentence adds value; no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of permission checking and no output schema, the description covers the essential outcomes (signed verdicts), binding mechanism, and verification path. It could go further by explaining what to do with the signed fact or edge cases, but the current information is sufficiently complete for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage with descriptions for all parameters. The tool description only summarizes the grant fields (allowed action, merchant/domain, etc.) without adding meaning beyond what the schema already documents. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks whether an agent's proposed action stays inside its declared permission grant, enumerating the grant dimensions (allowed action, merchant/domain, spend cap, budgets, etc.). It distinguishes itself from siblings by emphasizing it returns signed conformance facts, not judgments on intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes this is a mechanical conformance check and explicitly contrasts with 'a verdict on intent,' giving a sense of when to use it. It also points to a related tool (onyx_attestation_verify) for verification, but does not explicitly name alternative time-to-use scenarios or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_perm_grantAInspect
Mint a signed permission grant (onyx-perm-grant/v0) for an agent: a portable, Ed25519-notarized declaration of the scope it may act within (allowed_actions, allowed_merchants/domains, spend_max_usdc, principal, consent_ref, expires_at). Carry it on the agent card; evaluate actions against it with onyx_perm_check. Facts, not judgments — attests what was declared, not that the grantor holds the authority. (price: $0.00 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Who is granted (did:pkh / wallet / agent id) | |
| purpose | No | ||
| principal | No | Who grants the authority (did / email / org) | |
| expires_at | No | Unix time the grant lapses | |
| consent_ref | No | Proof of consent (AP2 mandate id / signature hash) | |
| velocity_max | No | Max number of transacts per window | |
| spend_max_usdc | No | Hard ceiling per action | |
| allowed_actions | No | Subset of ['read', 'verify', 'transact', 'sign', 'negotiate', 'subscribe']; deny-by-default | |
| allowed_domains | No | ||
| spend_window_sec | No | Window length in seconds (default 86400) | |
| allowed_merchants | No | ||
| spend_window_usdc | No | Optional rolling-window aggregate cap | |
| per_counterparty_budget | No | Max aggregate spend per single counterparty per window |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and provides valuable behavioral context: it is a portable, Ed25519-notarized declaration, and it explicitly states 'Facts, not judgments — attests what was declared, not that the grantor holds the authority.' This non-obvious limitation is critical for an agent to understand. It also discloses price/tier. It could mention more about authentication or side effects, but the disclosed traits go beyond a basic action statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized and front-loaded with the primary action. It packs essential facts (format, scope fields, verification pointer, non-judgment caveat, price) into three sentences. The parenthetical price block is a bit cluttered but acceptable. It could be trimmed, but it earns its place overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with no output schema and no annotations, the description covers the core purpose, the grant format, key scope fields, the portable nature, the verification path, and the critical epistemic limitation. It is sufficiently complete for an agent to understand when and how to invoke the tool, though it doesn't describe return values (which is acceptable given no output schema).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 77% (10 of 13 params), so the schema already explains most parameters. The description merely lists some of the fields (allowed_actions, allowed_merchants/domains, spend_max_usdc, principal, consent_ref, expires_at) without adding meaning beyond the schema's own descriptions. There is no extra semantic detail, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Mint a signed permission grant') and the resource ('onyx-perm-grant/v0 for an agent'), and differentiates it from the sibling onyx_perm_check by explicitly directing verification to that tool. The verb+resource+scope is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it mints a grant to be carried on the agent card, and explicitly points to onyx_perm_check for evaluating actions against it. While it doesn't list explicit when-not-to-use scenarios or other alternatives, the named sibling provides sufficient differentiation for an agent to choose correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_preflightAInspect
Preflight safety check for an x402-gated endpoint, BEFORE you pay it. Probes the URL live (~8s timeout), confirms it actually speaks x402 (a proper HTTP 402 with a parseable payment-requirements body — not a decoy demanding payment in prose), parses the advertised price to human USDC, sanity-checks the payTo address (0x format + EIP-55 checksum) and network (mainnet vs silently-testnet), and flags dead/cold/zombie endpoints and absurd price traps. Returns one signed verdict — OK, WARN, or AVOID — plus every disclosed flag behind it. Sign facts, not judgments: each flag traces to a checkable rule, not an opaque trust score. (price: $0.02 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The x402-gated endpoint to preflight (http:// or https://). | |
| timeout_seconds | No | Probe timeout in seconds. Clamped to [2, 15]. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and excels. It discloses the live probe, ~8s timeout, expected HTTP 402 response with parseable payment body, price conversion to USDC, address format/checksum checks, network mainnet vs testnet, flagging of dead/cold/zombie endpoints and price traps, and the signed verdict format. It even explains the 'sign facts, not judgments' principle. No annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense and front-loaded with the main purpose in the first clause. Each subsequent detail (checks, return format, pricing) serves a functional role. The 'Sign facts, not judgments' phrase is a bit stylistic but adds transparency. It earns its place, though somewhat verbose compared to the leanest examples.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple checks, output verdict, safety-critical), the description is remarkably complete. It explains the return value (OK/WARN/AVOID plus flags), covers edge cases (decoy, dead endpoints, price traps), includes pricing, and clarifies the verification principles. No output schema exists, so the description must explain return values, which it does.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description mentions the timeout ('~8s timeout') which aligns with the timeout_seconds parameter and default, but does not add substantial meaning beyond the schema. The url parameter is self-evident. No additional syntax or format details are provided that would elevate above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Preflight safety check for an x402-gated endpoint, BEFORE you pay it,' which is a specific verb ('preflight'), a clear resource (x402-gated endpoint), and a clear scope (safety check before payment). It distinguishes from siblings by focusing on x402 endpoints and payment preflight, not generic transactions or payments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context: use this tool before paying an x402 endpoint. It explicitly says 'BEFORE you pay it,' which tells the agent when to invoke it. It does not name alternative tools or exclusions, but the x402-specific scope and the timing are sufficient guidance for this context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_research_intelAInspect
Research intel — has someone solved X already? Queries 240M+ academic works via OpenAlex (includes arXiv preprints, conference papers, journal articles), ranks by citation count + recency + relevance, returns top N papers with one-line abstract excerpts, citation counts, and author names. Built for autonomous agents that need to check prior art before burning cycles re-deriving a known result. Fallback to Semantic Scholar. (price: $0.05 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Research question or keyword string. Plain English works; OpenAlex handles tokenization. | |
| top_n | No | How many papers to return. | |
| sort_by | No | Ranking. citations = highest cited first; recency = newest first; relevance = OpenAlex semantic match. | relevance |
| year_from | No | Optional: only return papers from this year onward. | |
| min_citations | No | Filter out papers with fewer than this many citations. Use 50+ to surface only well-known work. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the burden of disclosure. It reveals the data source (OpenAlex), scope (240M+ works, arXiv/preprints/conference/journal), ranking logic (citations + recency + relevance), return content (titles, abstracts, citation counts, authors), and a fallback to Semantic Scholar. It also discloses pricing. It doesn't mention rate limits or error handling, but overall it is quite transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose, uses three concise sentences plus a pricing note. Each sentence earns its place: purpose, use case, fallback, and cost. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (5 params, no output schema), the description tells the agent what it returns (top N papers with abstracts, citations, authors), where data comes from, how results are ranked, and when to use it. This is sufficient for an agent to invoke the tool correctly without needing an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers 100% of parameters with detailed descriptions, so the baseline is 3. The description adds minimal extra parameter context, only reinforcing that ranking uses citations/relevance/recency and that min_citations with 50+ surfaces well-known work. It doesn't provide syntax or formatting details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Research intel — has someone solved X already?' which immediately conveys the tool's purpose: searching academic literature for prior art. It specifies the verb 'Queries' and the resource '240M+ academic works via OpenAlex,' and clearly distinguishes itself from the sibling tools (payments, attestations, registry, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states it is 'Built for autonomous agents that need to check prior art before burning cycles re-deriving a known result,' providing a clear when-to-use scenario. It does not explicitly mention when not to use it, but the sibling tools are so distinct that no exclusions are necessary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_retail_price_checkAInspect
Ground-truth retail oracle. Give a product URL; get the real current price, currency, and in-stock state as actually fetched now — with the extraction source (JSON-LD / OpenGraph / microdata) as evidence. Covers the long tail of no-API shops where agents otherwise hallucinate prices. Never guesses: returns price=None with confidence='none' when the page exposes no machine-readable price. Use before an agent quotes, compares, or transacts on a price it would otherwise invent. (price: $0.02 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Full product page URL (http/https). The exact page whose price + availability you want observed. | |
| expect_price | No | Optional. A price you believe is current. If given, the result includes matches_expected:bool + drift so a caller can detect a stale/hallucinated quote. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that results are 'actually fetched now', that it returns price=None with confidence='none' when no machine-readable price exists, and that extraction source is provided as evidence. It also exposes metering cost, which is a useful disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise for the information conveyed, using a few punchy sentences. The pricing note at the end is extra but relevant. Front-loaded with 'Ground-truth retail oracle'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the return values (price, currency, in-stock, source, confidence), failure behavior, and use case. Without an output schema, this is adequate. It doesn't cover error cases but they are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description does not add specific parameter semantics beyond what the schema already provides (e.g., expect_price is described only in schema).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a ground-truth price oracle for retail product URLs, specifying it fetches current price, currency, in-stock state, and extraction source. It distinguishes from siblings like market_rank or merchant_fact_check by emphasizing real-time fetch and evidence-based extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use: 'Use before an agent quotes, compares, or transacts on a price it would otherwise invent.' It also clarifies it covers no-API shops and never guesses, implying limits. However, it doesn't name alternative tools or explicitly say when not to use, so slightly below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_secure_paymentAInspect
Secure-transaction RAIL: one signed clearance before an agent sends funds. Give recipient + amount (and optionally a contract address or counterparty ERC-8004 agent id); Onyx runs the full security stack — recipient firewall, contract audit, counterparty reputation — and returns a single PASS / REVIEW / FAIL verdict + risk score, plus the Onyx take-rate quote (bps of value secured). Ed25519-signed so the clearance is provable. The check a serious agent runs before moving real money. Onyx never takes custody. (price: $0.25 USDC, tier: premium)
| Name | Required | Description | Default |
|---|---|---|---|
| recipient | Yes | 0x recipient address the agent is about to pay (Base). | |
| amount_usdc | Yes | Amount about to be sent, in USDC. Drives both the risk threshold and the take-rate quote. | |
| contract_address | No | Optional. If the payment interacts with a contract, its 0x address — triggers a full contract audit. | |
| counterparty_agent_id | No | Optional. If paying another AI agent, its ERC-8004 id — triggers a reputation check. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and does an excellent job. It discloses the signed verdict (Ed25519), non-custody behavior, output format (PASS/REVIEW/FAIL + risk score + quote), and cost ($0.25 USDC). These are substantive behavioral traits beyond basic read/write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then efficiently packs behavioral details, output, safety, and pricing. The markety line 'The check a serious agent runs before moving real money' adds a bit of positioning but no functional loss. Overall concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers inputs, what the tool does, the output (verdict + risk score + quote), and important traits (signed, non-custody, cost). Since there is no output schema, it reasonably describes the return. It stops short of detailing the exact risk-scale or error conditions, but is complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter has a meaningful description. The tool description essentially restates the required/optional nature ('recipient + amount' and optional contract/agent id) without adding new semantic detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb+resource+outcome: 'Secure-transaction RAIL: one signed clearance before an agent sends funds.' It distinguishes itself from sibling tools by emphasizing it is the pre-payment security check that returns a verdict and signed clearance, not merely a transaction or audit tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: 'Give recipient + amount ... the check a serious agent runs before moving real money.' It implies exactly when to use (before sending funds), but does not explicitly name sibling alternatives or state when not to use, so it misses full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_security_flagsAInspect
Return the signed security posture of an agent: a list of OBSERVED, offline-verifiable security flags (insecure transport, no permission scope/principal/consent, no spend cap, unsigned identity, counterparty-blind, injection surface, unbounded skills). Pass an endpoint URL and/or the agent's record. Facts, not judgments — flags are conditions, never verdicts on intent. Ed25519-signed; verify free with onyx_attestation_verify. (price: $0.00 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| agent | No | Optional: the agent's record (name, description, skills, and any permission fields) for a deeper scan | |
| endpoint | No | The agent's endpoint URL |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full transparency burden. It discloses that the result is Ed25519-signed, offline-verifiable, based on observed conditions, and includes pricing/tier information. It does not mention error behaviors or what happens when both inputs are provided, but offers substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary purpose and remains dense with useful details. The parenthetical pricing/tier info is slightly extraneous, but every other sentence contributes to how the agent should use and interpret the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description compensates by listing the specific flag categories and explaining the signed, verifiable nature of the result. It stops short of describing the exact response shape or potential error states, but provides enough context for an agent to select and invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining the two parameters can be supplied together or alone ('endpoint URL and/or the agent's record'), which clarifies the optionality implied by zero required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Return the signed security posture of an agent' and enumerates the exact security flags delivered. It distinguishes itself from verification tools like onyx_attestation_verify by noting the output is signed and can be verified separately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly says to pass an endpoint URL and/or the agent's record, and clarifies the output should be interpreted as facts rather than judgments. It also points to onyx_attestation_verify for verification, but does not explicitly state when not to use this tool vs. siblings like onyx_agent_verify.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_signature_guardAInspect
Pre-signature firewall for OFF-CHAIN drains — the check before your agent signs an EIP-712 typed-data message (Permit, Permit2, Seaport order). These drain a wallet with no on-chain approval: the signature itself is the authorization. Give the typed-data; Onyx identifies what it authorizes, flags unlimited token-permit values, EOA/unverified spenders, NFT-order signatures, and bad deadlines, and returns a SIGNED ALLOW/REVIEW/BLOCK + plain-English explanation. Covers the #2 wallet-drain vector that on-chain tx checks miss entirely. (price: $0.10 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| typed_data | Yes | The full EIP-712 typed-data object the agent is about to sign: {domain, primaryType, types, message}. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the input behavior, what it flags, the output format (ALLOW/REVIEW/BLOCK + explanation), and the cost ($0.10 USDC). It lacks details on failure modes or rate limits, but for a read-only pre-signature check, it is highly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than two sentences but every clause serves a purpose: threat context, input, flags, output, and pricing. It is front-loaded with the core purpose and remains readable, though slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, it adequately explains return values (SIGNED ALLOW/REVIEW/BLOCK + explanation) and the risk flags. It covers the input contract, the reasoning, and cost. Minor gaps like error handling are not critical for this tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single typed_data parameter, and the schema description already explains its structure. The description adds context about what the tool does with the parameter but does not enrich the parameter semantics beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific action: checking EIP-712 typed-data before signing, and distinct resource: off-chain permit signatures. It differentiates from sibling tools like onyx_tx_guard by explicitly stating it covers the #2 wallet-drain vector that on-chain tx checks miss entirely, making its unique purpose obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use (before signing an EIP-712 message) and contrasts with on-chain tx checks, implying a preferred alternative for this vector. However, it does not name specific sibling tools or provide explicit when-not scenarios, though the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_token_riskAInspect
Signed token-security oracle. Give a token contract (and chain); get the real on-chain risk facts as read right now — honeypot status, buy/sell tax, mintable, ownership-reclaim, transfer-pausable, proxy, LP-lock, holder count — plus a transparent 0-100 risk score and verdict you can recompute from the itemized factors. Ed25519-signed + timestamped so an agent can PROVE it screened a token before buying. Use before any agent swaps into or quotes an ERC-20 it didn't issue. (price: $0.10 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| chain | No | Chain the token lives on. Name ('base','ethereum','bsc','polygon','arbitrum','optimism','avalanche') or numeric chain id. Default 'base'. | base |
| contract | Yes | Token contract address (0x… 20-byte hex) to screen. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that data is read live from-chain, that results are Ed25519-signed and timestamped, and that the risk score is recomputable from itemized factors. This goes beyond a basic description, though it omits potential caveats like failure modes or rate limits, so it's a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: it front-loads the tool's identity as a signed token-security oracle, then lists the risk facts, explains the signature/proof value, and closes with usage guidance and pricing. Every sentence carries useful content without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description provides a thorough picture: what inputs to give, what facts are returned, how the score is computed, why the signature matters, and when to use it. This is complete for a moderate-complexity tool with only two inputs and a rich output set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with both parameters already documented. The description adds minimal extra meaning—only restating that you give a contract and chain. It doesn't provide additional detail about address format or chain values beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool screens a token contract and returns on-chain risk facts, a risk score, and a verdict. It names specific outputs (honeypot status, taxes, mintable, etc.) and scope (given a token contract and chain). However, it does not explicitly differentiate from siblings like onyx_contract_audit or onyx_security_flags, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use before any agent swaps into or quotes an ERC-20 it didn't issue,' which gives clear when-to-use guidance. It does not mention when not to use it or alternatives, but this clear context earns a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_track_recordAInspect
Onyx's measured precision — the proof no other x402 security tool can show. Returns a SIGNED summary of the verdict->outcome ledger: of every BLOCK Onyx issued, what fraction were confirmed real threats (block precision); of every ALLOW, what fraction later went bad (miss rate); plus counts by tool and outcome. Built from real reported outcomes via onyx_outcome_report. Free and public — this is how you verify Onyx is calibrated, not just confident. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| tool | No | Optional. Restrict the track record to one tool (e.g. onyx_tx_guard). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the data source, that results are signed, the key metrics, and pricing/free tier. It doesn't mention authentication or return format details, but for a read-only metrics endpoint this is reasonable transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably compact but contains some marketing fluff ('the proof no other x402 security tool can show', 'not just confident') and redundant pricing mentions ('Free and public' plus '(price: $0 USDC, tier: free)'). The core technical content is front-loaded and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional parameter) and lack of output schema, the description covers the main return concepts and purpose. It doesn't specify exact JSON structure, but an agent can select and invoke it appropriately; slightly more detail on the signed envelope would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter (tool) is fully described in the schema, yielding 100% coverage, so the description adds minimal extra value. The mention of 'counts by tool and outcome' indirectly reinforces the optional filter, meeting the baseline for schema-covered parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Returns') and identifies the exact resource: a signed verdict->outcome ledger summary. It clearly differentiates from sibling security tools by emphasizing calibration verification, not detection or lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames when to use this tool: to verify Onyx's calibration, and states it is built from onyx_outcome_report. It doesn't name alternative tools for exclusion, but the context is clear enough that an agent would know it's for track-record/calibration questions rather than individual checks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_tx_guardAInspect
Pre-payment security firewall. Give the recipient address your agent is about to pay (Base); get a SIGNED ALLOW/REVIEW/BLOCK verdict + risk score from real on-chain checks: EOA-vs-contract, contract code/verification, account age (tx count), funding history, burn/null-address guard, and sink/honeypot heuristics. Catches paying a brand-new, unverified, or drain-shaped recipient BEFORE the money leaves. Never guesses — every field is observed on-chain and Ed25519-signed. (price: $0.10 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes | 0x recipient address your agent is about to send funds to (Base mainnet). | |
| amount_usdc | No | Optional amount about to be sent (USDC). Larger amounts raise the review threshold. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It is highly transparent: it discloses the output (signed verdict + risk score), the nature of checks (real on-chain, never guesses), the signing mechanism (Ed25519), and even pricing/tier. It does not detail potential error conditions or response structure, but the core behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact yet information-rich, leading with the high-level purpose, then detailing checks and usage. Every sentence contributes value, including the pricing note which helps an agent decide whether to invoke the tool. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity and the lack of an output schema, the description provides sufficient context: what the tool does, what input it needs, what output to expect (verdict + score), and an example of when to use it. It does not exhaustively detail output structure or edge cases, but it is complete enough for practical invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters described. The description adds meaningful context beyond the schema: it explains that 'address' is the recipient about to be paid on Base, and it clarifies that larger amounts raise the review threshold. This extra semantic value elevates it above the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a 'pre-payment security firewall' with a specific verb+resource: providing a recipient address and receiving a signed ALLOW/REVIEW/BLOCK verdict plus risk score. It enumerates the concrete on-chain checks performed, distinguishing it from sibling tools that focus on other aspects like payment gateways or general preflight checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies the usage context: 'before the money leaves' and 'pre-payment'. It clearly implies this tool is for vetting a recipient address prior to sending funds. However, it does not explicitly name alternatives or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_tx_preflightAInspect
Universal pre-sign firewall — the check before your agent signs ANY transaction. Give the tx (to, data, value); Onyx decodes the 4-byte selector + args, tells you in plain terms what it does, and flags the wallet-drain patterns: unlimited approve(), setApprovalForAll (blanket NFT approval), transfers to fresh/EOA recipients, raw ETH sends to unknown addresses, and calls to unverified targets. Returns a SIGNED ALLOW/REVIEW/BLOCK + human explanation. The single highest-frequency safety gate an on-chain agent has — every signed tx should pass through it. (price: $0.10 USDC, tier: metered)
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | The transaction's `to` address (target contract or recipient). Required. | |
| data | No | The transaction calldata (0x-hex). Omit/empty for a plain ETH transfer. | |
| value_wei | No | Optional. ETH value being sent, in wei (base-10 string). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly. It discloses specific behaviors: decoding selectors, flagging particular patterns (unlimited approve, setApprovalForAll, transfers to fresh/EOA, raw ETH sends, unverified targets), and returning a 'SIGNED ALLOW/REVIEW/BLOCK + human explanation.' It also states pricing and tiering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact yet dense, with each sentence earning its place. It fronts the core purpose, lists flags, specifies output, and ends with practical info (price/tier). No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately explains return values ('SIGNED ALLOW/REVIEW/BLOCK + human explanation') and main usage. It covers parameters and safety patterns. It falls short of covering edge cases or error handling, but for a pre-sign analysis tool the core is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning by explaining the parameter trio '(to, data, value)', noting that data can be omitted for plain ETH transfer, and that value is in wei. This goes slightly beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Universal pre-sign firewall' that 'decodes the 4-byte selector + args' and 'flags wallet-drain patterns.' It clearly distinguishes itself from siblings by positioning as the 'single highest-frequency safety gate' for any transaction before signing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use it 'before your agent signs ANY transaction' and that 'every signed tx should pass through it,' giving clear context. However, it does not mention excluded cases or alternative tools (e.g., deeper audits), so it lacks explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_verified_issueAInspect
Get your domain Onyx Verified. Onyx runs its published objective checks (live TLS, reachability, no off-domain redirect, registration age disclosed) and, on pass, issues an Ed25519-signed verified record onto the public Observation Log — instantly queryable at /merchant/{domain} — plus a live badge (served by Onyx, so it can't be faked or staled) and a machine-readable status URL agents check before they pay you. Valid 90 days, renewable. This attests your domain PASSED published checks, like a CA certificate — Onyx never claims you are 'honest' or 'safe'; the value is a neutral, verifiable, public presence agents can trust. (price: $2.00 USDC, tier: premium)
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | The domain to verify, e.g. shop.example.com (your storefront/agent endpoint). | |
| contact | No | Optional contact (email/handle) recorded with the issuance for renewal notices. | |
| agent_id | No | Optional agent id to cross-link this verified domain to an agent identity. |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the exact checks performed (live TLS, reachability, no off-domain redirect, registration age disclosed), the output artifacts (Observation Log record, live badge, status URL), the 90-day validity, the limitation that Onyx does not claim honesty/safety, and the badge-serving mechanism to prevent faking. This goes well beyond inferable behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is verbose and includes marketing-style phrasing and an analogy (like a CA certificate). While the core purpose is front-loaded, the text is a dense paragraph with multiple clauses. Some sentences could be tightened without losing important details, so it earns a 3 rather than higher.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the main outputs and context well, including what happens on pass, the queryable endpoint, the badge, and the status URL. However, there is no output schema and the description does not specify the exact function return value (e.g., JSON success object). Given the tool's side-effect nature, the coverage is strong but not fully complete, so a 4 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description does not add significant parameter-specific meaning beyond what the schema already provides for domain, contact, and agent_id. It mentions 'your domain' but does not elaborate on contact or agent_id usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: to get a domain Onyx Verified by running objective checks and issuing a signed record upon pass. It specifies the resource (domain) and the verb (verify/issue), and it distinguishes from sibling tools by focusing on domain verification issuance rather than verification of existing attestations or fact-checking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: domain owners use this to get verified so agents can trust them before payment. It mentions the use case and even the price/tier, but it does not explicitly name alternatives or exclusions (e.g., when not to use this vs. onyx_attestation_verify), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_verify_explainAInspect
Diagnose a failing x402 v2 /verify. Decodes a captured X-PAYMENT header, runs 10 rules (decode, schema, network/asset/payTo match, value sufficiency, EIP-3009 timing, signature shape, scheme) against expected paymentRequirements, and returns the FIRST failing rule with a plain-English fix. Catches the common case where the facilitator returns bare HTTP 402 (no body) because of JWT or schema fail upstream of the verifier. Stdlib-only, no install, no network. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| now_unix | No | Override current unix time for replay/CI use. Defaults to now. | |
| x_payment_b64 | No | Base64-encoded X-PAYMENT (v2 PAYMENT-SIGNATURE) header value. Optional if payment_payload provided. | |
| payment_payload | No | Decoded payment payload dict. Use this OR x_payment_b64. | |
| payment_requirements | Yes | Expected paymentRequirements from the 402 challenge ({scheme, network, payTo, asset, maxAmountRequired, maxTimeoutSeconds, ...}). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral disclosure burden. It reveals the internal process (decode header, run 10 rules), the output behavior (first failing rule with fix), and operational constraints (stdlib-only, no network, no install, zero cost). This goes well beyond a minimal statement and gives the agent confidence about side effects and limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense but well-structured paragraph. It front-loads the core action, lists the rules compactly, and includes the most critical edge case. The parenthetical about price/tier is minor noise but does not detract from the overall efficiency. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, input, process, output, and an edge case, which is comprehensive for a diagnostic tool with no output schema. It leaves a minor ambiguity about what happens when all 10 rules pass (no failing rule), but overall it gives an agent sufficient context to decide when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description enhances understanding by linking the X-PAYMENT header to x_payment_b64 and paymentRequirements to the rules, but it does not add extra per-parameter semantics beyond the schema. It neither compensates nor detracts, hence a solid 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific verb-resource pairing 'Diagnose a failing x402 v2 /verify', which precisely states what the tool does. It further distinguishes itself from sibling tools by detailing the rule-based diagnosis and the output of 'FIRST failing rule with a plain-English fix', making its purpose unambiguous and differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use the tool: to diagnose a failing /verify call, especially when a bare HTTP 402 is returned. It does not explicitly mention alternatives or when not to use it, but the scenario-driven guidance is clear and actionable, meriting a score above the 'implied only' level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onyx_x402_receipt_verifyAInspect
Verify an x402 USDC settlement on Base or Base Sepolia. Given a tx hash, decodes the USDC Transfer log and confirms (or refutes) a claim of the form: 'tx X moved $Y USDC from A to B'. Returns success status, actual decoded values, and a clear discrepancy report if any field doesn't match. Free tier — useful for agents reconciling spend and operators auditing inbound payments. (price: $0 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| network | No | Chain to query. Must match where the tx was mined. | base |
| tx_hash | Yes | 0x-prefixed 32-byte tx hash to verify. | |
| expected_to | No | Optional. Expected recipient address (0x...). If provided, verifier checks Transfer.to matches. | |
| expected_from | No | Optional. Expected sender address (0x...). If provided, verifier checks Transfer.from matches. | |
| expected_amount_usdc | No | Optional. Expected USDC amount (whole USDC, not atomic). If provided, verifier checks Transfer.value matches (within 0.000001 tolerance). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the core behavior: decodes the USDC Transfer log, confirms/refutes the claim, and returns success status, decoded values, and a discrepancy report. However, it does not mention what happens on invalid tx hashes, wrong network, or missing logs, nor does it explicitly confirm the operation is read-only. This is adequate but has clear gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with three sentences and a parenthetical. It front-loads the primary action and provides additional context on outputs and use cases. The pricing note is slightly redundant with 'Free tier' but does not detract significantly. Overall, it is well-structured and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description compensates by stating return values (success status, actual decoded values, discrepancy report). It also covers the main use cases and network scope. However, it omits error semantics and does not clarify how to distinguish between a failed verification vs. an invalid transaction, which would be helpful. The description is largely complete for a read-only verification tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers 100% of parameters with descriptions, so the baseline is 3. The description reinforces the meaning of expected parameters by framing them within the claim ('tx X moved $Y USDC from A to B'), but it does not add new parameter-level details beyond the schema, such as the tolerance for amounts. Thus, it meets the baseline without exceeding it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: verify an x402 USDC settlement on Base or Base Sepolia by decoding the USDC Transfer log and confirming/refuting a claim. It uses a specific verb ('verify') and resource ('x402 USDC settlement'), and it distinguishes itself from sibling tools like general transaction guards by focusing on x402 receipt verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: 'useful for agents reconciling spend and operators auditing inbound payments.' It implies when to use it (verifying a specific claim about a USDC transfer) but does not explicitly mention alternatives or when not to use it. Since it gives clear context without exclusions, it earns a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rhinogent_identityAInspect
Rhinogent identity card: give a self-custody Base (EVM) address, get back a signed, verifiable identity — deterministic callsign, did:pkh, and the agent's earned 0n1x-Verified credential. Keys never leave the agent; this issues only the public signed identity. The front door of the agent identity wallet. (price: $0.00 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes | The agent's self-custody Base/EVM address, 0x-prefixed (42 chars). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It significantly adds context: keys never leave the agent, only the public signed identity is issued, and the output is deterministic. This goes beyond schema and helps an agent understand safety and side-effect boundaries, though it does not mention reversibility or authorization requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with every sentence carrying meaningful information: what the tool does, key outputs, security behavior, and even pricing. The structure is efficient — no filler or redundancy — and the parenthetical pricing is a minor, relevant addition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description is remarkably complete. It covers the input type, the expected outputs, core behavioral guarantees (key custody, deterministic identity), and its role in the broader system. An agent has enough context to invoke it correctly without additional documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description repeats the 'self-custody Base (EVM)' aspect already present in the schema, but it does not add new semantic meaning about the address parameter beyond what the schema states. It aligns with the schema without enriching it further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: provide a self-custody Base address and receive a signed, verifiable identity. It enumerates specific outputs (deterministic callsign, did:pkh, 0n1x-Verified credential) and distinguishes itself from siblings by positioning as 'the front door of the agent identity wallet.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through phrases like 'front door of the agent identity wallet' and 'give a self-custody address, get back...', but it does not explicitly state when to use this tool versus alternatives like rhinogent_verify_counterparty or onyx_attestation_verify. There is no exclusion or alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rhinogent_mandateBInspect
Author a signed spend mandate for an agent: per-action and rolling-window USDC caps, merchant/domain allowlists, velocity limit and expiry — issued as a signed PERM_v0 grant the agent adopts. 'Transacts autonomously, within the boundaries you define.' (price: $0.00 USDC, tier: free)
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent the mandate authorizes (address or callsign). | |
| purpose | No | Human-readable purpose of the mandate. | |
| principal | No | Optional: the human/entity granting authority. | |
| expires_at | No | Unix time the mandate lapses. | |
| velocity_max | No | Max number of transactions per window. | |
| daily_cap_usdc | No | Rolling 24h aggregate cap (USDC). | |
| spend_max_usdc | No | Hard ceiling per single action (USDC). | |
| allowed_actions | No | Permitted action verbs (deny by default). | |
| allowed_domains | No | Domain allowlist (deny by default). | |
| allowed_merchants | No | Merchant allowlist (deny by default). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool issues a signed grant that the agent adopts, and that transactions are autonomous within defined limits. However, it does not mention side effects, irreversibility, or required permissions beyond the schema fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core description is a single dense sentence that conveys most needed info, but it is padded with a tagline and pricing info ('$0.00 USDC, tier: free') that do not aid tool selection. This is unnecessary fluff, though not overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 10 parameters, no output schema, and no annotations. The description explains the purpose and main constraints but does not clearly state the return value (the signed grant is implied but not explicit) or any prerequisites/preconditions. The schema covers parameter semantics, but behavioral and output details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning by grouping parameters into functional categories: 'per-action and rolling-window USDC caps' clarifies spend_max_usdc vs daily_cap_usdc, and 'velocity limit and expiry' highlights velocity_max and expires_at. This goes beyond the individual schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Author a signed spend mandate for an agent' with specific controls (caps, allowlists, velocity, expiry). It distinguishes this from generic permission tools by referencing 'PERM_v0 grant' and 'agent adopts', but does not explicitly contrast with siblings like onyx_perm_grant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when you need to set up autonomous agent spending with boundaries. It gives a clear context ('Transacts autonomously, within the boundaries you define') but does not mention when not to use it or explicitly name alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rhinogent_verify_counterpartyAInspect
Know-your-counterparty for agents. Before your agent pays a merchant, get Ed25519-signed raw facts about who it's paying: domain registration age, live TLS cert, reachability + off-domain redirects, brand-impersonation similarity, and observed price vs expected. Facts only (never 'legit/scam') — the one check every other agent wallet skips. (price: $0.25 USDC, tier: premium)
| Name | Required | Description | Default |
|---|---|---|---|
| brand | No | Optional brand the storefront claims to be (enables the impersonation-similarity observation). | |
| domain | Yes | Counterparty/storefront domain or URL the agent is about to pay. | |
| expected_price | No | Optional price the agent expects to pay (enables the price-deviation observation). |
Tool Definition Quality
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully discloses behavior. It states that output is Ed25519-signed raw facts (ensuring authenticity), enumerates the fact categories, and explicitly declares 'Facts only (never 'legit/scam')' to set expectations about the nature of the output. It also reveals pricing and tier, which are important operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, fitting key details into three sentences. It front-loads the core purpose. However, the marketing clause 'the one check every other agent wallet skips' does not directly aid tool selection or invocation, slightly diminishing the value of that sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple fact types), lack of output schema, and no annotations, the description adequately explains what the agent should expect: raw facts with Ed25519 signing. It lists the specific fact categories but does not detail error cases or the exact output format, leaving minor gaps that could affect agent workflows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning by linking the 'brand' parameter to brand-impersonation similarity and 'expected_price' to price-deviation observation, which clarifies the purpose of optional parameters beyond the schema's generic descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to obtain Ed25519-signed raw facts about a counterparty before payment. It specifies the resource (counterparty/storefront domain) and the concrete data points returned (domain registration age, TLS cert, reachability, redirects, brand similarity, price deviation). This strongly distinguishes it from sibling tools like onyx_merchant_fact_check by emphasizing pre-payment verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use this tool: 'Before your agent pays a merchant.' This is clear context. However, it does not explicitly mention alternatives or when not to use it, though the line 'the one check every other agent wallet skips' implies uniqueness. It could be improved by naming sibling tools to avoid confusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Claim this connector by publishing a /.well-known/glama.json file on your server's domain with the following structure:
{
"$schema": "https://glama.ai/mcp/schemas/connector.json",
"maintainers": [{ "email": "your-email@example.com" }]
}The email address must match the email associated with your Glama account. Once published, Glama will automatically detect and verify the file within a few minutes.
Control your server's listing on Glama, including description and metadata
Access analytics and receive server usage reports
Get monitoring and health status updates for your server
Feature your server to boost visibility and reach more users
For users:
Full audit trail – every tool call is logged with inputs and outputs for compliance and debugging
Granular tool control – enable or disable individual tools per connector to limit what your AI agents can do
Centralized credential management – store and rotate API keys and OAuth tokens in one place
Change alerts – get notified when a connector changes its schema, adds or removes tools, or updates tool definitions, so nothing breaks silently
For server owners:
Proven adoption – public usage metrics on your listing show real-world traction and build trust with prospective users
Tool-level analytics – see which tools are being used most, helping you prioritize development and documentation
Direct user feedback – users can report issues and suggest improvements through the listing, giving you a channel you would not have otherwise
The connector status is unhealthy when Glama is unable to successfully connect to the server. This can happen for several reasons:
The server is experiencing an outage
The URL of the server is wrong
Credentials required to access the server are missing or invalid
If you are the owner of this MCP connector and would like to make modifications to the listing, including providing test credentials for accessing the server, please contact support@glama.ai.
Discussions
No comments yet. Be the first to start the discussion!
Related MCP Servers
- AlicenseAqualityCmaintenancePay-per-call x402 data products on Base mainnet — sanctions screening, aviation weather, mortgage rates, US property dossier, title chain, wallet balance, and agent session auth. Every call settles in USDC with an on-chain receipt, no accounts or API keys.795MIT
- FlicenseAquality-maintenanceTrust infrastructure for AI agents on Base. DEX Spread Oracle (live Uniswap V3 prices), on-chain escrow, insurance pool, and collective knowledge base. 7 smart contracts. Pay-per-query via x402 micropayments in USDC.6

AfaAgent x402 API Suiteofficial
Flicense-qualityCmaintenance43 x402-enabled API tools — DeFi, wallet security, AI/ML, developer tools, SEO. Pay-per-call USDC on Base via x402 protocol.- AlicenseAqualityBmaintenanceAI consensus market oracle for crypto traders and autonomous agents. BUY/SELL/HOLD signals with 11-signal consensus (RSI, MACD, funding rate, Fear & Greed, congressional trading, Polymarket edges). Ed25519-signed. x402 micropayments on Base.91MIT
Your Connectors
Sign in to create a connector for this server.