Batcave Detective Basic
Server Details
Evidence-bound second-opinion audit of an agent conclusion against caller-supplied evidence.
- Status
- Healthy
- Uptime
- 39.3% over 23 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 5 tools
Each tool has a broadly distinct function: claim audit, decision review, JSON repair, Python execution, and multi-role deliberation. However, 'detective_basic' and 'evidence_judge' are both paid review/audit products that could be confused without reading their descriptions closely.
All names are snake_case, but there is no consistent structural pattern. Some are compound nouns (detective_basic, scenario_council, json_doctor), one is noun+verb (evidence_judge), and one is verb+noun (run_python), making the set feel ad hoc.
Five tools is a well-scoped count for a product-oriented server. Each tool occupies a clear niche—two paid review products and three utility/runtime capabilities—without redundancy or unnecessary bloat.
The set covers the advertised paid products and utility functions, including one tool for each major capability. A minor gap is the lack of an explicit product catalog or a combined workflow tool, and all tools only return checkout contracts rather than executing, though that appears intentional.
Available Tools
5 toolsdetective_basicDetective BasicBInspect
Batcave paid product at /v1/audit-claims. The canonical product contract is published by STORE OpenAPI. Price: 0.02 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| claims | Yes | ||
| job_id | No | ||
| evidence | No | ||
| importance | No | ||
| original_task | Yes | ||
| reasoning_summary | No | ||
| current_conclusion | Yes | ||
| freshness_expectation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden and does useful work: it states that this is a paid product, prices it at 0.02 USD via x402, and explicitly discloses that calling the tool does not execute the product or charge the caller. Details like auth and rate limits are left to the referenced OpenAPI contract, but the most important side-effect behavior is made clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the resource and behavior, then add cost and non-execution caveats. Every sentence adds information: endpoint/product, canonical contract reference, price, and what the call does and does not do.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has nine parameters, no output schema, and no annotations, but the description omits parameter semantics, input examples, expected response format, and when to prefer this over siblings. The reference to the canonical STORE OpenAPI contract helps, but an agent still lacks enough context to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate by explaining any of the nine parameters such as original_task, current_conclusion, claims, evidence, or importance. The endpoint path is the only hint about what the parameters relate to, which is insufficient for an agent to populate them confidently.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource (/v1/audit-claims, the Batcave detective_basic product), a specific action (returns the canonical REST/x402 checkout contract), and explicitly contrasts that with executing the product. It does not name or differentiate sibling tools, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: to obtain the checkout contract rather than to run the product, reinforced by 'does not execute the product or charge the caller yet.' However, it never explicitly states when not to use it or how it compares with evidence_judge, json_doctor, run_python, or scenario_council.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evidence_judgeEvidence JudgeCInspect
Batcave paid product at /v1/review-decision. The canonical product contract is published by STORE OpenAPI. Price: 0.02 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| risk | No | ||
| task | No | ||
| job_id | No | ||
| evidence | No | ||
| proposed_action | Yes | ||
| reasoning_summary | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden and does useful work: it reveals that this call only returns a checkout contract, is priced at $0.02 via x402, and triggers no charge or execution yet. Missing details like expected response shape or authorization are not disclosed, but the key side-effect boundary is clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying a distinct fact: endpoint/product, canonical contract source, price, and side-effect boundary. It is efficient, though the first sentence is somewhat telegraphic and assumes knowledge of 'Batcave'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter paid tool with no output schema and no annotations, the description is incomplete: it fails to map the input schema to the product contract, explain what the returned contract contains, or describe how to proceed after obtaining it. The side-effect guarantee is helpful but not sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description explains none of the six parameters (proposed_action, evidence, risk, task, job_id, reasoning_summary). An agent cannot tell what values to pass or how they affect the returned contract.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly names the endpoint (/v1/review-decision), identifies the tool as a paid Batcave product, and states the immediate behavior: it returns the canonical REST/x402 checkout contract rather than executing the product. This is a concrete verb+resource, though it does not tie back to the evidence-judging semantics or distinguish itself from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides no explicit when-to-use or alternative routing; it never mentions detective_basic or scenario_council. It does clarify a boundary (does not charge/execute), but an agent still has to infer when retrieving the checkout contract is the right step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
json_doctorJSON DoctorAInspect
Deterministic CPU-only utility. Repairs a bounded set of common JSON syntax mistakes, returns canonical JSON, and optionally validates a documented JSON Schema subset. Missing fields are reported, never invented. Price: 0.003 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | repair_and_validate | |
| text | Yes | JSON or JSON-like text to parse, conservatively repair, and return. | |
| job_id | No | ||
| schema | No | Optional structural JSON Schema. JSON Doctor validates the documented supported keyword subset and reports unsupported keywords explicitly. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the burden, and it excels: it discloses determinism, CPU-only execution, bounded repair scope, honesty about missing fields (never invented), price via x402, and the crucial side effect that calling returns a checkout contract without executing or charging the caller. This goes far beyond a basic tool summary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with the core purpose front-loaded. The first sentence states what the tool does; the second adds price and side-effect context. Every sentence earns its place, and there is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four parameters, a nested schema object, and no output schema, the description gives essential behavior but leaves gaps: it does not explain the mode enum values, the role of job_id, or what the REST/x402 checkout contract contains. The safety and cost context is strong, but the operational completeness is only partial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% (text and schema have descriptions; mode and job_id do not). The description adds some parameter-level meaning by noting optional validation (mapping to mode) and the bounded/subset nature of schema validation, but it does not clarify mode enum values or the purpose of job_id. It neither fully compensates for the gap nor leaves it entirely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool repairs a bounded set of common JSON syntax mistakes, returns canonical JSON, and optionally validates a JSON Schema subset. This specific verb-resource mapping distinguishes it from unrelated siblings like run_python or scenario_council.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for use: it is a deterministic CPU-only JSON repair/validation utility. It does not explicitly name alternatives or exclusions, but the sibling tools are unrelated, so the context is sufficient without further routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_pythonRun PythonCInspect
Execute supplied Python files in a fresh disposable CPU-only container. No outbound network, no GPU, and no host filesystem access beyond the disposable workspace. Returns bounded execution evidence, requested artifacts, and Execution Receipt v1. Price: 0.01 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| cpus | No | ||
| files | Yes | Relative file path -> UTF-8 text content. | |
| stdin | No | ||
| job_id | No | ||
| command | No | Optional argv; defaults to ['python','main.py']. | |
| runtime | No | ||
| memory_mb | No | ||
| artifact_paths | No | ||
| timeout_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description discloses key behavior: disposable CPU-only container, no network/GPU/host filesystem, price, and the critical two-phase fact that this call returns a checkout contract and does not execute/charge yet. The conflicting 'Execute' opening keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only three sentences and front-loads environment constraints and pricing. However, the contradictory execution vs checkout statements make the structure confusing rather than clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nine-parameter tool with no output schema and low schema coverage, the description omits how to complete the x402 checkout, what 'bounded execution evidence' contains, parameter details, and error behavior. The two-phase behavior is mentioned, but an agent cannot confidently predict the tool's return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 22%, and the description does not compensate for the nine parameters. It only hints at 'supplied Python files' and 'requested artifacts,' leaving cpus, memory_mb, timeout_seconds, stdin, job_id, and command semantics to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence says 'Execute supplied Python files...', but the last says 'it does not execute the product or charge the caller yet' and that the call returns a checkout contract. This internal contradiction makes the tool's actual purpose ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool rather than the listed siblings (detective_basic, evidence_judge, json_doctor, scenario_council) or what triggers checkout vs execution. Only mechanics and pricing are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scenario_councilScenario CouncilAInspect
Runs exactly four fixed roles (Operator, Skeptic, Buyer, Adversary) for 1-3 bounded rounds over only the caller-supplied question, context, and evidence. Returns every role position, evidence lineage, structured peer challenges, belief changes, disagreements, falsifiers, deterministic process-integrity results, and one independent Detective Basic audit of the four final conclusions. Does not retrieve outside facts, execute actions, treat peer statements as evidence, or treat majority agreement as truth. Price: 0.25 USD via x402. Calling this MCP tool returns the canonical REST/x402 checkout contract; it does not execute the product or charge the caller yet.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | ||
| rounds | No | ||
| context | No | ||
| evidence | No | ||
| question | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so thoroughly. It discloses that calling returns a REST/x402 checkout contract, does not execute the product or charge the caller, does not retrieve outside facts, does not execute actions, and does not treat majority agreement as truth. It also surfaces deterministic process-integrity results and a price.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: action, return payload, exclusions, pricing, and checkout side-effect. The most important scoping information is front-loaded in the first sentence, and the rest is tightly organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations and no output schema, the description covers action, inputs, output contents, constraints, cost, and side-effect behavior. An agent can understand what will happen, what it will receive, and what it will not be charged before invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for question, context, evidence, and rounds (via '1-3 bounded rounds'), but job_id is never mentioned and no parameter formats or constraints are described. Partial compensation with a clear gap for one optional parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb and resource: 'Runs exactly four fixed roles (Operator, Skeptic, Buyer, Adversary) for 1-3 bounded rounds...' and clearly scopes inputs to caller-supplied question, context, and evidence. It further differentiates itself from execution-style siblings by stating it does not retrieve outside facts or execute actions, and returns a Detective Basic audit rather than being the detective tool itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes clear context: use this tool for bounded multi-role analysis over caller-supplied material only, and it explicitly says it does not fetch outside facts or execute actions. It stops short of naming sibling tools or giving explicit when-to-use/when-not-to-use routing, so it misses the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
- First observed
detective_basic - First observed
evidence_judge - First observed
json_doctor - First observed
run_python - First observed
scenario_council
Related MCP Connectors
Second opinion before an irreversible agent action; signed proofs, free verify, public ledger.
Evidence-backed x402 web verification for AI agents, with auditable decisions for every condition.
Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations.
Independent agent preflight: verify change, mandates, evidence, and exact execution.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables agents to verify claims with evidence-based truth scores and confidence levels by running a deterministic pipeline of evidence lanes and adversarial checks.26MIT
- FlicenseNot gradedqualityNot gradedmaintenanceA bounded evidence review engine that ingests documents, extracts evidence for a given claim, detects contradictions, and produces auditable evidence packets without hallucinations or open-web research.-
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to verify claims against evidence through MCP tools, returning approved, rejected, or needs_review verdicts with replayable receipts.MIT
- AlicenseBqualityCmaintenanceEnables AI agents to verify technical claims against supplied evidence, identify unsupported assumptions and contradictions, and recommend the smallest next check before acting.5MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.