Skip to main content
Glama

run_sandbox_payment_probe

After fresh founder confirmation of the exact USD 1.00 Moov sandbox charge, create or replay one server-priced diagnostic order and send its server-configured card-payment source through Corply's durable create-and-authorize pipeline. The tool accepts no provider IDs, card/bank data, customer data, amount, currency, merchant, fee, reserve, or payout destination. It requires an active route, a fresh zero-variance reconciliation, server-held probe configuration, and an exact idempotency key. This is test-mode money movement only; it cannot charge live money, enable production, or prove settlement by itself. Prerequisite: authenticated active organization access plus every prerequisite stated above. Canonicality: invokes the shared backend action; trust the returned actual_tool_output and context_engineering instead of adding a state-recovery call. Idempotency: obey the tool-specific retry key or guarantee; if none is stated, inspect refreshed state before retrying. Confirmation boundary: no additional confirmation is needed for this read, reversible save, explicit fact/evidence record, link preparation, plan refresh, or action pre-authorized by a standing founder-configured policy.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
routeIdYes
companyIdNo
confirmationYes
idempotencyKeyYes
_corply_contextNoEcho context_engineering.context_session from the prior Corply result.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the minimal annotations by explaining test-mode-only behavior, the durable create-and-authorize pipeline, inability to prove settlement, canonicality expectations, idempotency retry guidance, and the fact that no additional confirmation is needed after pre-authorization. This is substantial behavioral context beyond readOnlyHint/destructiveHint/openWorldHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening is front-loaded with the action and prerequisites, but the description contains a large amount of generic boilerplate, especially the confirmation-boundary sentence listing unrelated categories like 'read, reversible save, explicit fact/evidence record, link preparation, plan refresh.' Several statements repeat the same point about prerequisites and confirmation, so not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, state-changing tool with no output schema, the description is quite complete: it covers prerequisites, behavior, limitations, canonicality, idempotency, and confirmation policy. The main gaps are lack of explicit companyId semantics, no mention of what the actual_tool_output contains, and no direct differentiation from the payout probe sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, so the description carries most of the burden for parameter meaning. It usefully explains routeId as an active route, idempotencyKey as an exact key, and confirmation as the founder-confirmed sandbox charge, but companyId is left unexplained and _corply_context relies on the schema description. Core required parameters are reasonably covered, but optional parameter semantics are incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: create or replay a server-priced diagnostic order and send the card-payment source through the durable create-and-authorize pipeline in Moov sandbox. It also disambiguates from production payment tools by emphasizing test-mode-only money movement. However, it does not explicitly contrast with the sibling run_sandbox_payout_probe, so some sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong when-to-use guidance: after fresh founder confirmation of the exact USD 1.00 charge, with an active route, fresh zero-variance reconciliation, server-held probe configuration, and an exact idempotency key. It also states a clear boundary: it cannot charge live money, enable production, or prove settlement. It does not explicitly name alternatives such as run_sandbox_payout_probe or when not to use this tool in favor of another.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.4/5.0
Disambiguation3/5

There is notable overlap among status/read tools (get_company_briefing, get_status, get_org, whoami) and a dense cluster of payment-related tools (request_payment, await_payment, create_payment_project, create_payment_route_draft, etc.). The detailed descriptions help differentiate them, but agents could still misselect when the surface is this large.

Naming Consistency4/5

Tool names overwhelmingly follow a verb_noun snake_case pattern (e.g., create_payment_route_draft, record_operating_event, start_bank_onboarding). Minor exceptions like 'whoami', 'recall', and 'remember' are acceptable single-verb commands, so the naming is highly consistent overall.

Tool Count2/5

At 52 tools, the server far exceeds the 25+ threshold for 'too many'. While the breadth of domains (formation, payments, cap table, operating compliance) somewhat justifies the count, it still feels heavy and likely increases selection errors and cognitive load for agents.

Completeness4/5

The tool surface covers the full formation lifecycle (save, validate, generate, sign, submit), payment handling, bank onboarding, cap table management, and operating records/evidence workflows. Minor gaps exist (e.g., no update/delete for existing companies, no explicit company dissolution), but the core workflows are well covered.