Skip to main content
Glama

Server Details

Agent cost, latency and quality analysis, report delivery and order status. Crypto payments paused.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Uptime
99.9% over 22 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A4.1/5.0

Scored across 10 tools

Disambiguation4/5

Most tools target clearly distinct functions: latency analysis, cost auditing, quality checks, catalog discovery, order lookup, and three named paid products. The main ambiguity is the three purchase_* tools, which share an identical payment flow and differ only by product name/price, but the product names (changes, evidence, snapshot) are explicit enough to separate them.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun style: analyze_agent_latency, audit_agent_costs, check_agent_quality, get_catalog, get_order, purchase_*, save_project_report. The pattern is predictable, with purchase tools sharing the same prefix and differing only by product.

Tool Count5/5

Ten tools is well-scoped for a paid agent-performance API: three analytical tools, three purchasable products, three supporting lookups/order/sample tools, and one report-saving tool. Each tool serves a distinct role without excessive redundancy or an excessively thin surface.

Completeness4/5

The surface covers the core workflow well: discover APIs, analyze performance, purchase three products, retrieve existing orders, and save reports. Minor gaps exist, such as no order-listing or report-reading capability, and get_free_sample feels peripheral to agent performance, but agents can still complete main tasks without dead ends.

Available Tools

10 tools
analyze_agent_latencyLatency LabA
Read-onlyIdempotent
Inspect

Find slow agent workflows and retry overhead before scaling. Supply recorded attempts with task_id, workflow, variant, cost_usd, success and optional latency_ms. Returns P50/P95 of summed recorded attempt durations per task, retry counts, duration coverage and an optional P95 threshold check for each workflow and variant. Missing durations produce null percentiles, not zero latency. This measures supplied durations, not live end-to-end service latency. Free calculation; requires an active ALPNAI agent key; no payment or automatic deployment.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsYes
configNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
groupsYes
measurementYes
schema_versionYes
percentile_methodYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, the description adds critical behavior: missing durations yield null percentiles rather than zero, the calculation is free and requires only an active key, and no payment or deployment occurs. It also clarifies that it analyzes supplied durations, not live service latency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: purpose, input shape, output summary, edge-case behavior, and operational constraints. It is front-loaded and compact despite covering many important caveats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema and output schema, the description fully equips the agent to understand input expectations, key output metrics, missing-data semantics, authentication needs, and non-deployment behavior. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With top-level schema description coverage at 0%, the description partially compensates by listing the required attempt fields and optional latency_ms. However, it gives only a vague nod to the config object ('optional P95 threshold check') and does not explain minSamples, minSuccessRate, monthlyTasks, or maxSuccessRateDrop.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific, actionable purpose: 'Find slow agent workflows and retry overhead before scaling.' It clearly identifies the resource (recorded agent attempts) and differentiates itself from the cost/quality-focused sibling tools by focusing on latency and retry counts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a clear use case ('before scaling') and an explicit exclusion ('not live end-to-end service latency'). It stops short of naming alternatives like audit_agent_costs or check_agent_quality, so some routing judgment is still left to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_agent_costsSpend Proof cost analysisB
Read-onlyIdempotent
Inspect

Free comparison of cost per successful task on supplied paired traces. Requires a sandbox agent key. No purchase, storage, external model call or automatic deployment. Success labels are supplied by the client.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsYes
configNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
configYes
purposeYes
workflowsYes
methodologyYes
schema_versionYes
data_provenanceYes
input_attempt_recordsYes
experimental_bias_controlYes
automatic_deployment_authorizedYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and idempotent, and the description adds valuable context: it explicitly states no purchase, storage, external model call, or automatic deployment, and that success labels are client-supplied. It also notes the sandbox key requirement, which is an operational constraint. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four short sentences with no fluff. The main purpose is front-loaded, and the caveats are listed efficiently. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested runs array with many fields and a config object), the description is minimal. It doesn't explain the expected input structure beyond 'paired traces', nor does it clarify the baseline/candidate semantics that the variant field implies. However, the output schema exists, so return values are covered. The description covers key constraints like the sandbox key and no external calls, but could benefit from a note on how to structure runs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention any parameters, and schema description coverage is 0%, so it fails to compensate for the schema's lack of explanatory text. The schema itself has some descriptions for fields like task_id and cost_usd, but the tool description adds no meaning. The 'paired traces' phrase hints at the variant field, but it's insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (compare) and resource (cost per successful task) on paired traces, making the tool's function clear. It does not explicitly name sibling alternatives, but the purpose is unambiguous. The phrase 'paired traces' hints at the baseline/candidate input structure.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus analyze_agent_latency or check_agent_quality. It only mentions a prerequisite (sandbox agent key) but not selection criteria. This is a significant gap for an agent deciding between analysis tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_agent_qualityQuality GateA
Read-onlyIdempotent
Inspect

Check whether a candidate agent workflow regresses before replacing the baseline. Supply baseline and candidate attempts on matching task IDs with client-provided success labels and costs. Returns observed success rates, task-set matching, sample-size, success, latency and cost gates, plus a decision such as collect_more_data, quality_regression or candidate_for_controlled_trial. Optional thresholds use config. It evaluates the recorded labels, not the correctness of answers or future performance. Free calculation; requires an active ALPNAI agent key; no payment or automatic deployment.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsYes
configNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
configYes
workflowsYes
schema_versionYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool read-only, idempotent, and non-destructive. The description adds substantial behavioral context beyond that: it evaluates recorded labels rather than correctness or future performance, has no payment or automatic deployment, requires an active key, and is free. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the main purpose, then flows from input, to output, to config, to caveats and access requirements. Every sentence carries distinct, necessary information with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex comparison tool with an output schema, the description covers input semantics, decision output types, optional config, key caveats, and authentication/payment constraints. Schema-level details like body limits and validation rules are already present, so the description does not need to repeat them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description gives semantic roles to both parameters: 'runs' are baseline/candidate attempts with client-provided success labels and costs, and 'config' supplies optional thresholds. This compensates for the low schema coverage, though it does not enumerate every config field or required field within runs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb ('check'), the resource ('candidate agent workflow'), and the decision frame ('before replacing the baseline'). It also clarifies the tool's quality focus, which distinguishes it from latency and cost analysis siblings. This is unambiguous and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when to use it—before replacing a baseline—and what inputs to supply: baseline and candidate attempts on matching task IDs. It does not explicitly name sibling alternatives or when-not conditions, but the context is strong enough for an agent to select the tool appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_catalogALPNAI service catalogA
Read-onlyIdempotent
Inspect

Discover cost, latency and quality analysis APIs, their inputs, outputs, free reproducible example, authentication and current payment availability. No payment or budget debit.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as read-only, idempotent, and non-destructive. The description adds the meaningful behavioral guarantee of 'No payment or budget debit' and notes the catalog includes authentication and payment availability, which helps an agent avoid expecting a paid or side-effecting operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the catalog contents are front-loaded and the clarifying no-debit statement is placed after the main payload. Every clause adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple no-parameter catalog lookup with no output schema, the description adequately lists what will be returned (API metadata, inputs/outputs, example, auth, payment availability) and the non-side-effect nature. It does not state the exact response format, but that is not critical for a zero-input introspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the input schema fully describes the call shape and the description adds no parameter burden. Baseline 4 applies because there is nothing for the description to explain about inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Discover') and identifies the resource (cost, latency, and quality analysis APIs) plus the exact metadata returned (inputs, outputs, example, authentication, payment availability). 'No payment or budget debit' clearly separates this catalog/introspection tool from the purchase_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this tool to explore available analysis APIs and their metadata, not to perform actions. 'No payment or budget debit' functions as an exclusion, but it does not explicitly name alternative tools such as purchase_changes or get_free_sample.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_free_sampleOpenAI source sampleA
Read-onlyIdempotent
Inspect

Read a dated primary-source sample about OpenAI. Not real-time or exhaustive.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that the sample is 'dated' and 'not real-time or exhaustive,' which is useful behavioral context beyond the annotations. It does not describe return format or pagination, but with zero parameters and a simple read operation, this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no wasted words. The core action and resource are front-loaded, and the caveat about not being real-time or exhaustive is placed immediately after. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool with no output schema, the description is nearly complete. It tells the agent what the sample is, its temporal nature, and its limitations. It could name a sibling for contrast, but nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no schema burden to compensate for. The description's mention of 'dated primary-source sample' gives the agent enough context to understand what the fixed output represents. Baseline 4 is appropriate for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read') and resource ('dated primary-source sample about OpenAI'), and clarifies it is not real-time or exhaustive. It is clear enough to distinguish from siblings like get_catalog or get_order, though it does not explicitly name a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when a dated primary-source sample about OpenAI is needed, and it warns against expecting real-time or exhaustive data. However, it does not explicitly state when to prefer an alternative sibling tool or provide exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_orderALPNAI order status and deliveryA
Idempotent
Inspect

Follow one existing commercial order owned by the authenticated agent, including after commerce is paused. May record a finalized blockchain receipt or release an expired, never-submitted reservation. Does not create an order, authorize spending or submit a payment. Pending results provide the same order ID and Retry-After; repeat get_order, never purchase again. Requires only the existing agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesThe original order_id returned by the purchase.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already indicate readOnlyHint=false and idempotentHint=true, and the description adds meaningful behavioral nuance: it may record a finalized blockchain receipt or release an expired, never-submitted reservation. It also discloses pending-response behavior and explicitly negates payment/order-creation side effects, going well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, each earning its place: purpose/scope, side effects and exclusions, then retry behavior and auth. No redundancy, no filler, and the most important scoping constraint is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter status tool with no output schema, the description is remarkably complete: it covers scope, ownership, side effects, what it does not do, pending/retry behavior, and authentication requirements. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter order_id is fully documented in the schema with a pattern and description ('The original order_id returned by the purchase'). The description adds context about reusing the same order ID for retries, but the schema already carries the essential parameter meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Follow') and resource ('one existing commercial order owned by the authenticated agent'), and the title clarifies 'status and delivery'. It clearly distinguishes itself from purchase-related sibling tools by explicitly saying it does not create an order, authorize spending, or submit a payment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use context ('including after commerce is paused') and clear retry guidance: pending results provide the same order ID and Retry-After, so the agent should repeat get_order and never purchase again. It also states the only prerequisite: the existing agent key.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

purchase_changesChange Set purchaseA
DestructiveIdempotent
Inspect

Get Change Set for 0.05 USDC. Default mode sandbox creates a test receipt and debits only the test budget. Explicit mode live uses x402 only when commerce, owner mandate and billing eligibility are enabled. An unpaid result contains PaymentRequired in structuredContent and content. Retry the same tool and idempotency key with the signed PaymentPayload in params._meta["x402/payment"]. Confirmed settlement is returned in result._meta["x402/payment-response"]. For a pending order, call get_order with the original order_id; do not submit another payment. REST continuation and PAYMENT-SIGNATURE headers remain supported. Never send a private key.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoSandbox is the default. Live explicitly requests the existing authorized x402 purchase flow.sandbox
sinceNo
mandate_idNoOwner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers.
idempotency_keyYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by describing the cost, test budget debit, PaymentRequired response, retry mechanics with signed PaymentPayload, settlement metadata, and REST header support. The security note about never sending a private key adds important operational context. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Information-dense and front-loaded with the core purchase action, followed by modes, retry flow, alternatives, and security. Every sentence contributes, though the dense paragraph format could be slightly easier to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the payment flow, retries, response fields, pending-order behavior, headers, and safety guidance, which is strong for a payment tool without an output schema. The main gap is the unexplained 'since' parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning for idempotency_key (retry with same key) and mode (sandbox/live conditions), which the schema only partially covers. However, the 'since' parameter remains undocumented by both schema and description, so it does not fully compensate for the 50% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get Change Set') plus the cost ('0.05 USDC'), making the tool's function immediately clear. It also distinguishes itself from the free sample flow and from get_order for pending orders.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains when to use sandbox vs live mode, when to retry with the same idempotency key, and when to call get_order instead of submitting another payment. It gives clear when/when-not guidance and names an alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

purchase_evidenceEvidence Pack purchaseA
DestructiveIdempotent
Inspect

Get Evidence Pack for 0.25 USDC. Default mode sandbox creates a test receipt and debits only the test budget. Explicit mode live uses x402 only when commerce, owner mandate and billing eligibility are enabled. An unpaid result contains PaymentRequired in structuredContent and content. Retry the same tool and idempotency key with the signed PaymentPayload in params._meta["x402/payment"]. Confirmed settlement is returned in result._meta["x402/payment-response"]. For a pending order, call get_order with the original order_id; do not submit another payment. REST continuation and PAYMENT-SIGNATURE headers remain supported. Never send a private key.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoSandbox is the default. Live explicitly requests the existing authorized x402 purchase flow.sandbox
sinceNo
mandate_idNoOwner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers.
idempotency_keyYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the annotations by explaining what each mode does, where PaymentRequired appears, how to retry with params._meta['x402/payment'], and where settlement appears in result._meta. It also warns 'Never send a private key.' Nothing contradicts the readOnlyHint=false, idempotentHint=true, or destructiveHint=true annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the core purpose, and every major behavioral point earns its place. A few trailing details like 'REST continuation and PAYMENT-SIGNATURE headers remain supported' add minor noise, but the overall structure is efficient for a payment-flow tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description compensates well by describing PaymentRequired, retry flow, settlement response location, and pending-order fallback. The main gap is the undocumented 'since' parameter, which leaves a small ambiguity in an otherwise complete operational picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, so the description must compensate. It adds real meaning for idempotency_key by tying it to PaymentPayload retries, and for mode by explaining sandbox versus live eligibility. However, the 'since' parameter remains completely unexplained in both schema and description, preventing a top score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Get Evidence Pack for 0.25 USDC.' It also clarifies the sandbox versus live behavior, making the tool's purpose unmistakable and distinct from nearby retrieval tools like get_order and get_catalog.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use conditions, including 'Explicit mode live uses x402 only when commerce, owner mandate and billing eligibility are enabled,' and it names the alternative: 'For a pending order, call get_order with the original order_id; do not submit another payment.' This is strong, actionable routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

purchase_snapshotSnapshot purchaseA
DestructiveIdempotent
Inspect

Get Snapshot for 0.01 USDC. Default mode sandbox creates a test receipt and debits only the test budget. Explicit mode live uses x402 only when commerce, owner mandate and billing eligibility are enabled. An unpaid result contains PaymentRequired in structuredContent and content. Retry the same tool and idempotency key with the signed PaymentPayload in params._meta["x402/payment"]. Confirmed settlement is returned in result._meta["x402/payment-response"]. For a pending order, call get_order with the original order_id; do not submit another payment. REST continuation and PAYMENT-SIGNATURE headers remain supported. Never send a private key.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoSandbox is the default. Live explicitly requests the existing authorized x402 purchase flow.sandbox
sinceNo
mandate_idNoOwner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers.
idempotency_keyYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Although annotations indicate readOnlyHint=false and destructiveHint=true, the description adds substantial behavioral detail: test receipts vs. live x402 flow, unpaid results containing PaymentRequired, how to retry with PaymentPayload, where settlement appears in the result, and the warning never to send a private key. This is far more than the annotations alone convey and does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the core purpose, then logically covers modes, failure handling, and retry behavior. It is longer than minimal but every sentence contributes meaningful operational context, though some details like 'REST continuation' could be trimmed without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a payment flow, no output schema, and only four parameters, the description is highly complete. It covers error states, idempotency, retry semantics, settlement metadata, and the alternative get_order path. The only notable omission is the meaning of the 'since' parameter, which prevents a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%, and the description compensates well for mode, idempotency_key, and mandate-related eligibility. However, the 'since' parameter remains completely unexplained in both the schema and the description, leaving a gap for an agent trying to understand what value to provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: purchasing a Snapshot for 0.01 USDC, and distinguishes between sandbox and live modes. This is a specific verb+resource with enough detail to identify it among siblings such as purchase_changes, purchase_evidence, and get_free_sample.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance: sandbox is default, live only applies when specific conditions are met, an unpaid result should be retried with the same tool and idempotency key, and pending orders should use get_order instead of submitting another payment. It names the alternative tool (get_order) and conveys a clear when-to-use/when-not-to-use distinction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_project_reportSave an agent performance reportA
Idempotent
Inspect

Compute and save an aggregate performance report in the fixed Projects destination authorized by the account owner. Uses the existing free or paid report allowance. Requires an active agent key and explicit owner write permission. Reuse request_id with identical content to retry within the same permission grant. No report reading, deletion, subscription or payment authorization. Send measurements without secrets.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesRecorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes.
titleYes
request_idYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide idempotentHint=true, readOnlyHint=false, and destructiveHint=false. The description adds valuable behavior beyond those: permission requirements, allowance consumption, fixed destination, retry semantics within a permission grant, and a security directive ('Send measurements without secrets'). There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences front-load the core purpose and destination, then pack the necessary permissions, idempotency, exclusions, and data-sensitivity warning. Every clause adds useful guidance and there is no filler or duplication of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the nested input schema, the description supplies essential invocation context: authorization, destination, allowance, retry, and 'no reading/deletion/payment' boundaries. The main gap is that there is no output schema and the description does not state what the tool returns or how failure manifests, which would make the calling contract fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, and the description compensates somewhat by explaining request_id retry semantics and warning that input should contain measurements, not secrets. However, it does not explain title's purpose, config semantics, or how the input object maps to report computation. The schema itself covers some parameters, but the description leaves meaning gaps for the remaining ones.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Compute and save an aggregate performance report in the fixed Projects destination.' This clearly states the action, object, and destination, and the explicit exclusions ('No report reading, deletion, subscription or payment authorization') distinguish it from sibling analysis and purchase tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: it requires an active agent key and explicit owner write permission, uses the report allowance, and supports idempotent retry via request_id. It states exclusions ('No report reading, deletion, subscription or payment authorization') but does not explicitly name sibling alternatives or state 'use X when Y' conditions, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Addedget_order
  2. 3 tool updates
    • Changedpurchase_changes2 fields changed
      • addedInput schema / properties / mandate_id
        Added value: +{
        +  "description": "Owner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers.",
        +  "pattern": "^[-a-zA-Z0-9]{8,80}$",
        +  "type": "string"
        +}
      • addedInput schema / properties / mode
        Added value: +{
        +  "default": "sandbox",
        +  "description": "Sandbox is the default. Live explicitly requests the existing authorized x402 purchase flow.",
        +  "enum": [
        +    "sandbox",
        +    "live"
        +  ],
        +  "type": "string"
        +}
    • Changedpurchase_evidence2 fields changed
      • addedInput schema / properties / mandate_id
        Added value: +{
        +  "description": "Owner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers.",
        +  "pattern": "^[-a-zA-Z0-9]{8,80}$",
        +  "type": "string"
        +}
      • addedInput schema / properties / mode
        Added value: +{
        +  "default": "sandbox",
        +  "description": "Sandbox is the default. Live explicitly requests the existing authorized x402 purchase flow.",
        +  "enum": [
        +    "sandbox",
        +    "live"
        +  ],
        +  "type": "string"
        +}
    • Changedpurchase_snapshot2 fields changed
      • addedInput schema / properties / mandate_id
        Added value: +{
        +  "description": "Owner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers.",
        +  "pattern": "^[-a-zA-Z0-9]{8,80}$",
        +  "type": "string"
        +}
      • addedInput schema / properties / mode
        Added value: +{
        +  "default": "sandbox",
        +  "description": "Sandbox is the default. Live explicitly requests the existing authorized x402 purchase flow.",
        +  "enum": [
        +    "sandbox",
        +    "live"
        +  ],
        +  "type": "string"
        +}
  3. 4 tool updates
    • Changedanalyze_agent_latency9 fields changed
      • addedInput schema / description
        Added value: +"Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
      • addedInput schema / properties / config / additionalProperties
        Added value: +false
      • addedInput schema / properties / config / properties
        Added value: +{
        +  "maxP95LatencyMs": {
        +    "maximum": 86400000000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "maxSuccessRateDrop": {
        +    "default": 0.02,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "minSamples": {
        +    "default": 30,
        +    "maximum": 500,
        +    "minimum": 2,
        +    "type": "integer"
        +  },
        +  "minSuccessRate": {
        +    "default": 0.95,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "monthlyTasks": {
        +    "description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied.",
        +    "maximum": 1000000,
        +    "minimum": 1,
        +    "type": "integer"
        +  }
        +}
      • addedInput schema / properties / config / required
        Added value: +[]
      • addedInput schema / properties / runs / items / additionalProperties
        Added value: +false
      • addedInput schema / properties / runs / items / properties
        Added value: +{
        +  "cost_usd": {
        +    "description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals.",
        +    "maximum": 10000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "latency_ms": {
        +    "description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured.",
        +    "maximum": 86400000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "success": {
        +    "type": "boolean"
        +  },
        +  "task_id": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units.",
        +    "maxLength": 128,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  },
        +  "variant": {
        +    "enum": [
        +      "baseline",
        +      "candidate"
        +    ],
        +    "type": "string"
        +  },
        +  "workflow": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units.",
        +    "maxLength": 80,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  }
        +}
      • addedInput schema / properties / runs / items / required
        Added value: +[
        +  "task_id",
        +  "workflow",
        +  "variant",
        +  "cost_usd",
        +  "success"
        +]
      • addedInput schema / title
        Added value: +"ALPNAI recorded agent attempts"
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "additionalProperties": true,
        +  "properties": {
        +    "groups": {
        +      "items": {
        +        "additionalProperties": true,
        +        "properties": {
        +          "duration_coverage": {
        +            "maximum": 1,
        +            "minimum": 0,
        +            "type": "number"
        +          },
        +          "max_ms": {
        +            "anyOf": [
        +              {
        +                "maximum": 86400000000,
        +                "minimum": 0,
        +                "type": "number"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "p50_ms": {
        +            "anyOf": [
        +              {
        +                "maximum": 86400000000,
        +                "minimum": 0,
        +                "type": "number"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "p95_ms": {
        +            "anyOf": [
        +              {
        +                "maximum": 86400000000,
        +                "minimum": 0,
        +                "type": "number"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "recorded_attempts": {
        +            "maximum": 1000,
        +            "minimum": 1,
        +            "type": "integer"
        +          },
        +          "retry_attempts": {
        +            "maximum": 999,
        +            "minimum": 0,
        +            "type": "integer"
        +          },
        +          "tasks": {
        +            "maximum": 1000,
        +            "minimum": 1,
        +            "type": "integer"
        +          },
        +          "threshold_ms": {
        +            "anyOf": [
        +              {
        +                "maximum": 86400000000,
        +                "minimum": 0,
        +                "type": "number"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "variant": {
        +            "enum": [
        +              "baseline",
        +              "candidate"
        +            ],
        +            "type": "string"
        +          },
        +          "within_p95_threshold": {
        +            "anyOf": [
        +              {
        +                "type": "boolean"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "workflow": {
        +            "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units.",
        +            "maxLength": 80,
        +            "minLength": 1,
        +            "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +            "type": "string"
        +          }
        +        },
        +        "required": [
        +          "workflow",
        +          "variant",
        +          "tasks",
        +          "recorded_attempts",
        +          "retry_attempts",
        +          "duration_coverage",
        +          "p50_ms",
        +          "p95_ms",
        +          "max_ms",
        +          "threshold_ms",
        +          "within_p95_threshold"
        +        ],
        +        "type": "object"
        +      },
        +      "maxItems": 1000,
        +      "minItems": 1,
        +      "type": "array"
        +    },
        +    "measurement": {
        +      "const": "summed_recorded_attempt_durations_per_task",
        +      "type": "string"
        +    },
        +    "percentile_method": {
        +      "const": "nearest_rank",
        +      "type": "string"
        +    },
        +    "schema_version": {
        +      "const": "1.0.0",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "schema_version",
        +    "measurement",
        +    "percentile_method",
        +    "groups"
        +  ],
        +  "type": "object"
        +}
    • Changedaudit_agent_costs9 fields changed
      • addedInput schema / description
        Added value: +"Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
      • addedInput schema / properties / config / additionalProperties
        Added value: +false
      • addedInput schema / properties / config / properties
        Added value: +{
        +  "maxP95LatencyMs": {
        +    "maximum": 86400000000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "maxSuccessRateDrop": {
        +    "default": 0.02,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "minSamples": {
        +    "default": 30,
        +    "maximum": 500,
        +    "minimum": 2,
        +    "type": "integer"
        +  },
        +  "minSuccessRate": {
        +    "default": 0.95,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "monthlyTasks": {
        +    "description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied.",
        +    "maximum": 1000000,
        +    "minimum": 1,
        +    "type": "integer"
        +  }
        +}
      • addedInput schema / properties / config / required
        Added value: +[]
      • addedInput schema / properties / runs / items / additionalProperties
        Added value: +false
      • addedInput schema / properties / runs / items / properties
        Added value: +{
        +  "cost_usd": {
        +    "description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals.",
        +    "maximum": 10000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "latency_ms": {
        +    "description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured.",
        +    "maximum": 86400000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "success": {
        +    "type": "boolean"
        +  },
        +  "task_id": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units.",
        +    "maxLength": 128,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  },
        +  "variant": {
        +    "enum": [
        +      "baseline",
        +      "candidate"
        +    ],
        +    "type": "string"
        +  },
        +  "workflow": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units.",
        +    "maxLength": 80,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  }
        +}
      • addedInput schema / properties / runs / items / required
        Added value: +[
        +  "task_id",
        +  "workflow",
        +  "variant",
        +  "cost_usd",
        +  "success"
        +]
      • addedInput schema / title
        Added value: +"ALPNAI recorded agent attempts"
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "additionalProperties": true,
        +  "properties": {
        +    "automatic_deployment_authorized": {
        +      "const": false,
        +      "type": "boolean"
        +    },
        +    "config": {
        +      "additionalProperties": true,
        +      "properties": {
        +        "maxP95LatencyMs": {
        +          "maximum": 86400000000,
        +          "minimum": 0,
        +          "type": "number"
        +        },
        +        "maxSuccessRateDrop": {
        +          "default": 0.02,
        +          "maximum": 1,
        +          "minimum": 0,
        +          "type": "number"
        +        },
        +        "minSamples": {
        +          "default": 30,
        +          "maximum": 500,
        +          "minimum": 2,
        +          "type": "integer"
        +        },
        +        "minSuccessRate": {
        +          "default": 0.95,
        +          "maximum": 1,
        +          "minimum": 0,
        +          "type": "number"
        +        },
        +        "monthlyTasks": {
        +          "description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied.",
        +          "maximum": 1000000,
        +          "minimum": 1,
        +          "type": "integer"
        +        }
        +      },
        +      "required": [
        +        "minSamples",
        +        "minSuccessRate",
        +        "maxSuccessRateDrop"
        +      ],
        +      "type": "object"
        +    },
        +    "data_provenance": {
        +      "const": "user_supplied_not_verified",
        +      "type": "string"
        +    },
        +    "experimental_bias_control": {
        +      "const": "unknown",
        +      "type": "string"
        +    },
        +    "input_attempt_records": {
        +      "maximum": 1000,
        +      "minimum": 1,
        +      "type": "integer"
        +    },
        +    "methodology": {
        +      "additionalProperties": true,
        +      "properties": {
        +        "duplicates": {
        +          "type": "string"
        +        },
        +        "latency": {
        +          "type": "string"
        +        },
        +        "logical_task_key": {
        +          "items": {
        +            "type": "string"
        +          },
        +          "maxItems": 3,
        +          "minItems": 3,
        +          "type": "array"
        +        },
        +        "missing_costs": {
        +          "type": "string"
        +        },
        +        "quality": {
        +          "type": "string"
        +        },
        +        "success": {
        +          "type": "string"
        +        },
        +        "uncertainty": {
        +          "type": "string"
        +        }
        +      },
        +      "required": [
        +        "logical_task_key",
        +        "success",
        +        "duplicates",
        +        "latency",
        +        "quality",
        +        "uncertainty",
        +        "missing_costs"
        +      ],
        +      "type": "object"
        +    },
        +    "purpose": {
        +      "const": "cost_per_successful_agent_task_audit",
        +      "type": "string"
        +    },
        +    "schema_version": {
        +      "const": "1.0.0",
        +      "type": "string"
        +    },
        +    "workflows": {
        +      "items": {
        +        "additionalProperties": true,
        +        "properties": {
        +          "automatic_deployment_authorized": {
        +            "const": false,
        +            "type": "boolean"
        +          },
        +          "baseline": {
        +            "anyOf": [
        +              {
        +                "additionalProperties": true,
        +                "properties": {
        +                  "attempts": {
        +                    "maximum": 1000,
        +                    "minimum": 1,
        +                    "type": "integer"
        +                  },
        +                  "cost_per_successful_task_usd": {
        +                    "anyOf": [
        +                      {
        +                        "minimum": 0,
        +                        "type": "number"
        +                      },
        +                      {
        +                        "type": "null"
        +                      }
        +                    ]
        +                  },
        +                  "cost_per_task_usd": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "latency_complete": {
        +                    "type": "boolean"
        +                  },
        +                  "p95_recorded_attempt_latency_per_task_ms": {
        +                    "anyOf": [
        +                      {
        +                        "maximum": 86400000000,
        +                        "minimum": 0,
        +                        "type": "number"
        +                      },
        +                      {
        +                        "type": "null"
        +                      }
        +                    ]
        +                  },
        +                  "retry_attempts": {
        +                    "maximum": 999,
        +                    "minimum": 0,
        +                    "type": "integer"
        +                  },
        +                  "success_rate": {
        +                    "maximum": 1,
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "success_rate_interval_95_wilson": {
        +                    "items": {
        +                      "maximum": 1,
        +                      "minimum": 0,
        +                      "type": "number"
        +                    },
        +                    "maxItems": 2,
        +                    "minItems": 2,
        +                    "type": "array"
        +                  },
        +                  "successful_tasks": {
        +                    "maximum": 1000,
        +                    "minimum": 0,
        +                    "type": "integer"
        +                  },
        +                  "tasks": {
        +                    "maximum": 1000,
        +                    "minimum": 1,
        +                    "type": "integer"
        +                  },
        +                  "total_cost_usd": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  }
        +                },
        +                "required": [
        +                  "tasks",
        +                  "attempts",
        +                  "retry_attempts",
        +                  "successful_tasks",
        +                  "total_cost_usd",
        +                  "cost_per_task_usd",
        +                  "cost_per_successful_task_usd",
        +                  "success_rate",
        +                  "success_rate_interval_95_wilson",
        +                  "p95_recorded_attempt_latency_per_task_ms",
        +                  "latency_complete"
        +                ],
        +                "type": "object"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "candidate": {
        +            "anyOf": [
        +              {
        +                "additionalProperties": true,
        +                "properties": {
        +                  "attempts": {
        +                    "maximum": 1000,
        +                    "minimum": 1,
        +                    "type": "integer"
        +                  },
        +                  "cost_per_successful_task_usd": {
        +                    "anyOf": [
        +                      {
        +                        "minimum": 0,
        +                        "type": "number"
        +                      },
        +                      {
        +                        "type": "null"
        +                      }
        +                    ]
        +                  },
        +                  "cost_per_task_usd": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "latency_complete": {
        +                    "type": "boolean"
        +                  },
        +                  "p95_recorded_attempt_latency_per_task_ms": {
        +                    "anyOf": [
        +                      {
        +                        "maximum": 86400000000,
        +                        "minimum": 0,
        +                        "type": "number"
        +                      },
        +                      {
        +                        "type": "null"
        +                      }
        +                    ]
        +                  },
        +                  "retry_attempts": {
        +                    "maximum": 999,
        +                    "minimum": 0,
        +                    "type": "integer"
        +                  },
        +                  "success_rate": {
        +                    "maximum": 1,
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "success_rate_interval_95_wilson": {
        +                    "items": {
        +                      "maximum": 1,
        +                      "minimum": 0,
        +                      "type": "number"
        +                    },
        +                    "maxItems": 2,
        +                    "minItems": 2,
        +                    "type": "array"
        +                  },
        +                  "successful_tasks": {
        +                    "maximum": 1000,
        +                    "minimum": 0,
        +                    "type": "integer"
        +                  },
        +                  "tasks": {
        +                    "maximum": 1000,
        +                    "minimum": 1,
        +                    "type": "integer"
        +                  },
        +                  "total_cost_usd": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  }
        +                },
        +                "required": [
        +                  "tasks",
        +                  "attempts",
        +                  "retry_attempts",
        +                  "successful_tasks",
        +                  "total_cost_usd",
        +                  "cost_per_task_usd",
        +                  "cost_per_successful_task_usd",
        +                  "success_rate",
        +                  "success_rate_interval_95_wilson",
        +                  "p95_recorded_attempt_latency_per_task_ms",
        +                  "latency_complete"
        +                ],
        +                "type": "object"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "comparison": {
        +            "additionalProperties": true,
        +            "properties": {
        +              "baseline_only_tasks": {
        +                "maximum": 1000,
        +                "minimum": 0,
        +                "type": "integer"
        +              },
        +              "candidate_only_tasks": {
        +                "maximum": 1000,
        +                "minimum": 0,
        +                "type": "integer"
        +              },
        +              "matched_task_ids": {
        +                "maximum": 1000,
        +                "minimum": 0,
        +                "type": "integer"
        +              },
        +              "same_task_set": {
        +                "type": "boolean"
        +              }
        +            },
        +            "required": [
        +              "matched_task_ids",
        +              "baseline_only_tasks",
        +              "candidate_only_tasks",
        +              "same_task_set"
        +            ],
        +            "type": "object"
        +          },
        +          "decision": {
        +            "enum": [
        +              "missing_comparison",
        +              "collect_more_data",
        +              "quality_regression",
        +              "latency_data_required",
        +              "latency_regression",
        +              "no_economic_advantage",
        +              "candidate_for_controlled_trial"
        +            ],
        +            "type": "string"
        +          },
        +          "gates": {
        +            "additionalProperties": true,
        +            "properties": {
        +              "both_variants": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "experimental_bias_control": {
        +                "const": "unknown",
        +                "type": "string"
        +              },
        +              "lower_cost_per_successful_task": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "minimum_distinct_tasks_per_variant": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "observed_success_rate": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "recorded_latency": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "same_task_set": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              }
        +            },
        +            "required": [
        +              "both_variants",
        +              "minimum_distinct_tasks_per_variant",
        +              "same_task_set",
        +              "observed_success_rate",
        +              "recorded_latency",
        +              "lower_cost_per_successful_task",
        +              "experimental_bias_control"
        +            ],
        +            "type": "object"
        +          },
        +          "opportunity_monthly": {
        +            "anyOf": [
        +              {
        +                "additionalProperties": true,
        +                "properties": {
        +                  "baseline_expected_cost_usd": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "baseline_launched_tasks_per_month": {
        +                    "maximum": 1000000,
        +                    "minimum": 1,
        +                    "type": "integer"
        +                  },
        +                  "basis": {
        +                    "const": "same_expected_successful_task_volume",
        +                    "type": "string"
        +                  },
        +                  "candidate_expected_cost_usd": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "candidate_expected_launched_tasks_per_month": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "conditions": {
        +                    "items": {
        +                      "type": "string"
        +                    },
        +                    "minItems": 1,
        +                    "type": "array"
        +                  },
        +                  "currency": {
        +                    "const": "USD",
        +                    "type": "string"
        +                  },
        +                  "excludes": {
        +                    "items": {
        +                      "type": "string"
        +                    },
        +                    "minItems": 1,
        +                    "type": "array"
        +                  },
        +                  "expected_successful_tasks_per_month": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "period": {
        +                    "const": "month",
        +                    "type": "string"
        +                  },
        +                  "potential_cost_difference_usd": {
        +                    "minimum": 0,
        +                    "type": "number"
        +                  },
        +                  "status": {
        +                    "const": "conditional_projection_not_realized",
        +                    "type": "string"
        +                  }
        +                },
        +                "required": [
        +                  "status",
        +                  "basis",
        +                  "currency",
        +                  "period",
        +                  "baseline_launched_tasks_per_month",
        +                  "expected_successful_tasks_per_month",
        +                  "candidate_expected_launched_tasks_per_month",
        +                  "baseline_expected_cost_usd",
        +                  "candidate_expected_cost_usd",
        +                  "potential_cost_difference_usd",
        +                  "excludes",
        +                  "conditions"
        +                ],
        +                "type": "object"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "realized_savings_usd": {
        +            "type": "null"
        +          },
        +          "workflow": {
        +            "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units.",
        +            "maxLength": 80,
        +            "minLength": 1,
        +            "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +            "type": "string"
        +          }
        +        },
        +        "required": [
        +          "workflow",
        +          "baseline",
        +          "candidate",
        +          "comparison",
        +          "gates",
        +          "decision",
        +          "automatic_deployment_authorized",
        +          "realized_savings_usd",
        +          "opportunity_monthly"
        +        ],
        +        "type": "object"
        +      },
        +      "maxItems": 1000,
        +      "minItems": 1,
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "schema_version",
        +    "purpose",
        +    "data_provenance",
        +    "config",
        +    "input_attempt_records",
        +    "experimental_bias_control",
        +    "automatic_deployment_authorized",
        +    "methodology",
        +    "workflows"
        +  ],
        +  "type": "object"
        +}
    • Changedcheck_agent_quality9 fields changed
      • addedInput schema / description
        Added value: +"Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
      • addedInput schema / properties / config / additionalProperties
        Added value: +false
      • addedInput schema / properties / config / properties
        Added value: +{
        +  "maxP95LatencyMs": {
        +    "maximum": 86400000000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "maxSuccessRateDrop": {
        +    "default": 0.02,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "minSamples": {
        +    "default": 30,
        +    "maximum": 500,
        +    "minimum": 2,
        +    "type": "integer"
        +  },
        +  "minSuccessRate": {
        +    "default": 0.95,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "monthlyTasks": {
        +    "description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied.",
        +    "maximum": 1000000,
        +    "minimum": 1,
        +    "type": "integer"
        +  }
        +}
      • addedInput schema / properties / config / required
        Added value: +[]
      • addedInput schema / properties / runs / items / additionalProperties
        Added value: +false
      • addedInput schema / properties / runs / items / properties
        Added value: +{
        +  "cost_usd": {
        +    "description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals.",
        +    "maximum": 10000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "latency_ms": {
        +    "description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured.",
        +    "maximum": 86400000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "success": {
        +    "type": "boolean"
        +  },
        +  "task_id": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units.",
        +    "maxLength": 128,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  },
        +  "variant": {
        +    "enum": [
        +      "baseline",
        +      "candidate"
        +    ],
        +    "type": "string"
        +  },
        +  "workflow": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units.",
        +    "maxLength": 80,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  }
        +}
      • addedInput schema / properties / runs / items / required
        Added value: +[
        +  "task_id",
        +  "workflow",
        +  "variant",
        +  "cost_usd",
        +  "success"
        +]
      • addedInput schema / title
        Added value: +"ALPNAI recorded agent attempts"
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "additionalProperties": true,
        +  "properties": {
        +    "config": {
        +      "additionalProperties": true,
        +      "properties": {
        +        "maxP95LatencyMs": {
        +          "maximum": 86400000000,
        +          "minimum": 0,
        +          "type": "number"
        +        },
        +        "maxSuccessRateDrop": {
        +          "default": 0.02,
        +          "maximum": 1,
        +          "minimum": 0,
        +          "type": "number"
        +        },
        +        "minSamples": {
        +          "default": 30,
        +          "maximum": 500,
        +          "minimum": 2,
        +          "type": "integer"
        +        },
        +        "minSuccessRate": {
        +          "default": 0.95,
        +          "maximum": 1,
        +          "minimum": 0,
        +          "type": "number"
        +        },
        +        "monthlyTasks": {
        +          "description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied.",
        +          "maximum": 1000000,
        +          "minimum": 1,
        +          "type": "integer"
        +        }
        +      },
        +      "required": [
        +        "minSamples",
        +        "minSuccessRate",
        +        "maxSuccessRateDrop"
        +      ],
        +      "type": "object"
        +    },
        +    "schema_version": {
        +      "const": "1.0.0",
        +      "type": "string"
        +    },
        +    "workflows": {
        +      "items": {
        +        "additionalProperties": true,
        +        "properties": {
        +          "automatic_deployment_authorized": {
        +            "const": false,
        +            "type": "boolean"
        +          },
        +          "baseline_success": {
        +            "anyOf": [
        +              {
        +                "maximum": 1,
        +                "minimum": 0,
        +                "type": "number"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "candidate_success": {
        +            "anyOf": [
        +              {
        +                "maximum": 1,
        +                "minimum": 0,
        +                "type": "number"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "comparison": {
        +            "additionalProperties": true,
        +            "properties": {
        +              "baseline_only_tasks": {
        +                "maximum": 1000,
        +                "minimum": 0,
        +                "type": "integer"
        +              },
        +              "candidate_only_tasks": {
        +                "maximum": 1000,
        +                "minimum": 0,
        +                "type": "integer"
        +              },
        +              "matched_task_ids": {
        +                "maximum": 1000,
        +                "minimum": 0,
        +                "type": "integer"
        +              },
        +              "same_task_set": {
        +                "type": "boolean"
        +              }
        +            },
        +            "required": [
        +              "matched_task_ids",
        +              "baseline_only_tasks",
        +              "candidate_only_tasks",
        +              "same_task_set"
        +            ],
        +            "type": "object"
        +          },
        +          "decision": {
        +            "enum": [
        +              "missing_comparison",
        +              "collect_more_data",
        +              "quality_regression",
        +              "latency_data_required",
        +              "latency_regression",
        +              "no_economic_advantage",
        +              "candidate_for_controlled_trial"
        +            ],
        +            "type": "string"
        +          },
        +          "gates": {
        +            "additionalProperties": true,
        +            "properties": {
        +              "both_variants": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "experimental_bias_control": {
        +                "const": "unknown",
        +                "type": "string"
        +              },
        +              "lower_cost_per_successful_task": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "minimum_distinct_tasks_per_variant": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "observed_success_rate": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "recorded_latency": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              },
        +              "same_task_set": {
        +                "enum": [
        +                  "pass",
        +                  "fail",
        +                  "unknown",
        +                  "not_requested"
        +                ],
        +                "type": "string"
        +              }
        +            },
        +            "required": [
        +              "both_variants",
        +              "minimum_distinct_tasks_per_variant",
        +              "same_task_set",
        +              "observed_success_rate",
        +              "recorded_latency",
        +              "lower_cost_per_successful_task",
        +              "experimental_bias_control"
        +            ],
        +            "type": "object"
        +          },
        +          "workflow": {
        +            "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units.",
        +            "maxLength": 80,
        +            "minLength": 1,
        +            "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +            "type": "string"
        +          }
        +        },
        +        "required": [
        +          "workflow",
        +          "comparison",
        +          "gates",
        +          "decision",
        +          "baseline_success",
        +          "candidate_success",
        +          "automatic_deployment_authorized"
        +        ],
        +        "type": "object"
        +      },
        +      "maxItems": 1000,
        +      "minItems": 1,
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "schema_version",
        +    "config",
        +    "workflows"
        +  ],
        +  "type": "object"
        +}
    • Changedsave_project_report8 fields changed
      • addedInput schema / properties / input / description
        Added value: +"Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
      • addedInput schema / properties / input / properties / config / additionalProperties
        Added value: +false
      • addedInput schema / properties / input / properties / config / properties
        Added value: +{
        +  "maxP95LatencyMs": {
        +    "maximum": 86400000000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "maxSuccessRateDrop": {
        +    "default": 0.02,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "minSamples": {
        +    "default": 30,
        +    "maximum": 500,
        +    "minimum": 2,
        +    "type": "integer"
        +  },
        +  "minSuccessRate": {
        +    "default": 0.95,
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "monthlyTasks": {
        +    "description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied.",
        +    "maximum": 1000000,
        +    "minimum": 1,
        +    "type": "integer"
        +  }
        +}
      • addedInput schema / properties / input / properties / config / required
        Added value: +[]
      • addedInput schema / properties / input / properties / runs / items / additionalProperties
        Added value: +false
      • addedInput schema / properties / input / properties / runs / items / properties
        Added value: +{
        +  "cost_usd": {
        +    "description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals.",
        +    "maximum": 10000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "latency_ms": {
        +    "description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured.",
        +    "maximum": 86400000,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  "success": {
        +    "type": "boolean"
        +  },
        +  "task_id": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units.",
        +    "maxLength": 128,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  },
        +  "variant": {
        +    "enum": [
        +      "baseline",
        +      "candidate"
        +    ],
        +    "type": "string"
        +  },
        +  "workflow": {
        +    "description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units.",
        +    "maxLength": 80,
        +    "minLength": 1,
        +    "pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
        +    "type": "string"
        +  }
        +}
      • addedInput schema / properties / input / properties / runs / items / required
        Added value: +[
        +  "task_id",
        +  "workflow",
        +  "variant",
        +  "cost_usd",
        +  "success"
        +]
      • addedInput schema / properties / input / title
        Added value: +"ALPNAI recorded agent attempts"
  4. 1 tool update
    • Addedsave_project_report
  5. 8 tool updates
    • First observedanalyze_agent_latency
    • First observedaudit_agent_costs
    • First observedcheck_agent_quality
    • First observedget_catalog
    • First observedget_free_sample
    • First observedpurchase_changes
    • First observedpurchase_evidence
    • First observedpurchase_snapshot

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources