Batcave Agent Services
Server Details
MCP discovery and x402 checkout handoff for five paid Batcave agent products.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 5 tools
Tools target distinct product endpoints, but detective_basic, evidence_judge, and scenario_council all relate to audit/review/reasoning, and scenario_council explicitly includes a Detective Basic audit. The uniform 'returns checkout contract' behavior further blurs operational differences, making selection less clear.
All names use snake_case, which is consistent. However, the semantic pattern is mixed: four are noun-based (detective_basic, evidence_judge, json_doctor, scenario_council) while run_python is verb_noun, a minor deviation from a predictable convention.
Five tools is well within the 3-15 range and each corresponds to a distinct product in the Batcave service catalog. The set is neither bloated nor too thin for the stated purpose.
The server exposes five specific products, but none of the tools execute the product or return results; they only return checkout contracts. There is no generic product discovery, status, or execution tool, leaving agents unable to complete a full service workflow via MCP alone.
Available Tools
5 toolsdetective_basicDetective BasicBInspect
Batcave paid product at /v1/audit-claims. The canonical product contract is published by STORE OpenAPI. Price: 0.02 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| claims | Yes | ||
| job_id | No | ||
| evidence | No | ||
| importance | No | ||
| original_task | Yes | ||
| reasoning_summary | No | ||
| current_conclusion | Yes | ||
| freshness_expectation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden and does useful work: it states that this is a paid product, prices it at 0.02 USD via x402, and explicitly discloses that calling the tool does not execute the product or charge the caller. Details like auth and rate limits are left to the referenced OpenAPI contract, but the most important side-effect behavior is made clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the resource and behavior, then add cost and non-execution caveats. Every sentence adds information: endpoint/product, canonical contract reference, price, and what the call does and does not do.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has nine parameters, no output schema, and no annotations, but the description omits parameter semantics, input examples, expected response format, and when to prefer this over siblings. The reference to the canonical STORE OpenAPI contract helps, but an agent still lacks enough context to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate by explaining any of the nine parameters such as original_task, current_conclusion, claims, evidence, or importance. The endpoint path is the only hint about what the parameters relate to, which is insufficient for an agent to populate them confidently.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource (/v1/audit-claims, the Batcave detective_basic product), a specific action (returns the canonical REST/x402 checkout contract), and explicitly contrasts that with executing the product. It does not name or differentiate sibling tools, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: to obtain the checkout contract rather than to run the product, reinforced by 'does not execute the product or charge the caller yet.' However, it never explicitly states when not to use it or how it compares with evidence_judge, json_doctor, run_python, or scenario_council.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evidence_judgeEvidence JudgeCInspect
Batcave paid product at /v1/review-decision. The canonical product contract is published by STORE OpenAPI. Price: 0.02 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| risk | No | ||
| task | No | ||
| job_id | No | ||
| evidence | No | ||
| proposed_action | Yes | ||
| reasoning_summary | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden and does useful work: it reveals that this call only returns a checkout contract, is priced at $0.02 via x402, and triggers no charge or execution yet. Missing details like expected response shape or authorization are not disclosed, but the key side-effect boundary is clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying a distinct fact: endpoint/product, canonical contract source, price, and side-effect boundary. It is efficient, though the first sentence is somewhat telegraphic and assumes knowledge of 'Batcave'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter paid tool with no output schema and no annotations, the description is incomplete: it fails to map the input schema to the product contract, explain what the returned contract contains, or describe how to proceed after obtaining it. The side-effect guarantee is helpful but not sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description explains none of the six parameters (proposed_action, evidence, risk, task, job_id, reasoning_summary). An agent cannot tell what values to pass or how they affect the returned contract.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly names the endpoint (/v1/review-decision), identifies the tool as a paid Batcave product, and states the immediate behavior: it returns the canonical REST/x402 checkout contract rather than executing the product. This is a concrete verb+resource, though it does not tie back to the evidence-judging semantics or distinguish itself from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides no explicit when-to-use or alternative routing; it never mentions detective_basic or scenario_council. It does clarify a boundary (does not charge/execute), but an agent still has to infer when retrieving the checkout contract is the right step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
json_doctorJSON DoctorAInspect
Deterministic CPU-only utility. Repairs a bounded set of common JSON syntax mistakes, returns canonical JSON, and optionally validates a documented JSON Schema subset. Missing fields are reported, never invented. Price: 0.003 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | repair_and_validate | |
| text | Yes | JSON or JSON-like text to parse, conservatively repair, and return. | |
| job_id | No | ||
| schema | No | Optional structural JSON Schema. JSON Doctor validates the documented supported keyword subset and reports unsupported keywords explicitly. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the burden, and it excels: it discloses determinism, CPU-only execution, bounded repair scope, honesty about missing fields (never invented), price via x402, and the crucial side effect that calling returns a checkout contract without executing or charging the caller. This goes far beyond a basic tool summary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with the core purpose front-loaded. The first sentence states what the tool does; the second adds price and side-effect context. Every sentence earns its place, and there is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four parameters, a nested schema object, and no output schema, the description gives essential behavior but leaves gaps: it does not explain the mode enum values, the role of job_id, or what the REST/x402 checkout contract contains. The safety and cost context is strong, but the operational completeness is only partial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% (text and schema have descriptions; mode and job_id do not). The description adds some parameter-level meaning by noting optional validation (mapping to mode) and the bounded/subset nature of schema validation, but it does not clarify mode enum values or the purpose of job_id. It neither fully compensates for the gap nor leaves it entirely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool repairs a bounded set of common JSON syntax mistakes, returns canonical JSON, and optionally validates a JSON Schema subset. This specific verb-resource mapping distinguishes it from unrelated siblings like run_python or scenario_council.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for use: it is a deterministic CPU-only JSON repair/validation utility. It does not explicitly name alternatives or exclusions, but the sibling tools are unrelated, so the context is sufficient without further routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_pythonRun PythonCInspect
Execute supplied Python files in a fresh disposable CPU-only container. No outbound network, no GPU, and no host filesystem access beyond the disposable workspace. Returns bounded execution evidence, requested artifacts, and Execution Receipt v1. Price: 0.01 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| cpus | No | ||
| files | Yes | Relative file path -> UTF-8 text content. | |
| stdin | No | ||
| job_id | No | ||
| command | No | Optional argv; defaults to ['python','main.py']. | |
| runtime | No | ||
| memory_mb | No | ||
| artifact_paths | No | ||
| timeout_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description discloses key behavior: disposable CPU-only container, no network/GPU/host filesystem, price, and the critical two-phase fact that this call returns a checkout contract and does not execute/charge yet. The conflicting 'Execute' opening keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only three sentences and front-loads environment constraints and pricing. However, the contradictory execution vs checkout statements make the structure confusing rather than clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nine-parameter tool with no output schema and low schema coverage, the description omits how to complete the x402 checkout, what 'bounded execution evidence' contains, parameter details, and error behavior. The two-phase behavior is mentioned, but an agent cannot confidently predict the tool's return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 22%, and the description does not compensate for the nine parameters. It only hints at 'supplied Python files' and 'requested artifacts,' leaving cpus, memory_mb, timeout_seconds, stdin, job_id, and command semantics to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence says 'Execute supplied Python files...', but the last says 'it does not execute the product or charge the caller yet' and that the call returns a checkout contract. This internal contradiction makes the tool's actual purpose ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool rather than the listed siblings (detective_basic, evidence_judge, json_doctor, scenario_council) or what triggers checkout vs execution. Only mechanics and pricing are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scenario_councilScenario CouncilAInspect
Compare competing interpretations or paths using only caller-supplied evidence. Runs exactly four fixed roles (Operator, Skeptic, Buyer, Adversary) for 1-3 bounded rounds over only the caller-supplied question, context, and evidence. Returns every role position, evidence lineage, structured peer challenges, belief changes, disagreements, falsifiers, deterministic process-integrity results, and one independent Detective Basic audit of the four final conclusions. Does not retrieve outside facts, execute actions, treat peer statements as evidence, or treat majority agreement as truth. Price: 0.25 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | ||
| rounds | No | ||
| context | No | ||
| evidence | No | ||
| question | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so: it discloses the fixed role set, round bounds, that outputs include evidence lineage, peer challenges, belief changes, falsifiers and a Detective Basic audit, and critically that calling the tool returns a REST/x402 checkout contract without executing the product or charging. That mutation/payment boundary is exactly the kind of disclosure an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: mechanism first, exclusions next, then the return-set enumeration and the payment/checkout caveat last. Long sentences, but nearly every clause adds a constraint or output fact, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema, no-annotation tool of this complexity, the description enumerates the return fields and states the execution/payment boundary, which is unusually complete. The remaining gap is the undocumented job_id and evidence-identifier semantics, which the description leaves entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description must compensate and only partly does. It clarifies that context and evidence are caller-supplied (with evidence presumably bounded) and that rounds are 1-3, but job_id is never explained and evidence item/id semantics are left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (compares competing interpretations/paths) and pins the mechanism precisely: four fixed named roles (Operator, Skeptic, Buyer, Adversary), 1-3 bounded rounds, over caller-supplied question/context/evidence only. An agent can distinguish this from siblings like evidence_judge or detective_basic without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong boundary conditions: only caller-supplied evidence, no outside retrieval, no action execution, no treating peer statements as evidence or majority agreement as truth. However it never explicitly routes the agent versus alternatives such as evidence_judge or when a single-role audit (detective_basic) is the better choice, so it stops just short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
- First observed
detective_basic - First observed
evidence_judge - First observed
json_doctor - First observed
run_python - First observed
scenario_council
Related MCP Connectors
Agent floor and MCP market: catalog, rack, Stripe checkout, x402, A2A.
Market-validated MCP capabilities with x402 paid execution.
MCP commerce surface for refurbished datacenter hardware, with x402 agent-payment discovery.
Workflow diagnostics, capability routing, and x402 settlement for MCP-compatible agents.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceAn MCP server that tests agent payment capabilities by making a real x402 payment ($10 USDC on Base) and returning a signed diagnostic report covering the entire payment flow.MIT
- FlicenseNot gradedqualityDmaintenancePaid remote MCP for hosted MCP server providing structured receipts, usage logs, and audit-ready evidence for agent and CI workflows.-
- AlicenseNot gradedqualityCmaintenanceEnables agents to access production-grade paid MCP tools with real on-chain x402 v2 settlement, including EVM wallet risk scoring, payload normalization, and facilitator discovery, all discoverable via Bazaar-compatible metadata.MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for AgentPay — the payment gateway for autonomous AI agents. Fund a wallet once, give your agent the key, and it discovers, provisions, and pays for tool APIs on its own. One key, every tool.112 npm1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.