Skip to main content
Glama

find_evidence

Fact-check or substantiate a claim against the corpus. Given a textual claim, retrieves and CLASSIFIES evidence into supporting / contradicting / neutral groups. Uses HyDE (hypothetical document expansion) — server generates plausible supporting/contradicting text, embeds, retrieves, then ranks by relation to original claim. Returns chunks with selfContained flag (safe-to-cite indicator). Use for fact-verification, controversy mapping, 'is this claim known?' queries. Modes: 'fast' (symmetric-by-construction grouping — returns grouped evidence but NO supporting/contradicting counts, because they would be symmetric by construction) / 'deep' (independent NLI classification; one LLM call per chunk, median 52 s measured). IMPORTANT: in 'fast' mode the supporting/contradicting counts are approximately balanced BY CONSTRUCTION and do NOT reflect actual literature distribution. Use 'deep' when measuring controversy balance, literature distribution, or any claim of the form 'the field is split N:M on this'.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNo'fast' (~3s): retrieval uses symmetric HyDE pools — top-20 chunks against the supporting-hypothetical plus top-20 against the contradicting-hypothetical, then each chunk is assigned to the bucket whose HyDE-vector it scored higher against. Because the retrieval pool is symmetric and the classification mirrors the retrieval direction, supporting/contradicting counts come out approximately balanced regardless of the actual distribution of evidence in the corpus (a topic that is 90% supported in the literature will still show a ~1:1 split here). Use fast mode for 'is there evidence on either side?', not for 'how is the field actually split?'. May also misclassify chunks that mention the topic but logically point the other way (e.g. a paper explaining 'BN is bad in transformers' may land in the contradicting bucket for an 'LN > BN' claim). 'deep': adds an independent per-chunk LLM NLI classification on top of the union pool, so counts reflect actual semantic distribution and can be arbitrarily asymmetric. ★ COST IS LINEAR IN POOL SIZE — one LLM call PER CHUNK, up to ~60 per invocation. Measured over 35 real calls: median 52 s end-to-end (the '~10s' this description used to claim was never the general case). That is money as well as time. Use 'deep' whenever classification accuracy or distribution shape matters — including controversy mapping and any analysis that interprets the supporting/contradicting ratio as a signal about the field.fast
claimYesStatement to fact-check or substantiate
limitNoMax results PER group (supporting/contradicting/neutral)
detailNo
run_idNoOptional. The active methodist run_id (as returned by the methodist diagnose / get_current_dose door). Pass it whenever you call this tool while working inside a run, so the call is attributed to that run for the §8 usage crosscheck — attribution is run-anchored, so it stays correct even if your access token refreshes mid-run. Must be YOUR run: a run_id owned by a different principal, or a non-existent run_id, is rejected.
categoriesNo
selfContainedOnlyNoIf true, only return chunks marked as understandable without prior context (safer to cite)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedInput schema / properties / mode / description
      Previous value: -"'fast' (~3s): retrieval uses symmetric HyDE pools — top-20 chunks against the supporting-hypothetical plus top-20 against the contradicting-hypothetical, then each chunk is assigned to the bucket whose HyDE-vector it scored higher against. Because the retrieval pool is symmetric and the classification mirrors the retrieval direction, supporting/contradicting counts come out approximately balanced regardless of the actual distribution of evidence in the corpus (a topic that is 90% supported in the literature will still show a ~1:1 split here). Use fast mode for 'is there evidence on either side?', not for 'how is the field actually split?'. May also misclassify chunks that mention the topic but logically point the other way (e.g. a paper explaining 'BN is bad in transformers' may land in the contradicting bucket for an 'LN > BN' claim). 'deep' (~10s): adds an independent per-chunk LLM NLI classification on top of the union pool, so counts reflect actual semantic distribution and can be arbitrarily asymmetric. Use 'deep' whenever classification accuracy or distribution shape matters — including controversy mapping and any analysis that interprets the supporting/contradicting ratio as a signal about the field."New value: +"'fast' (~3s): retrieval uses symmetric HyDE pools — top-20 chunks against the supporting-hypothetical plus top-20 against the contradicting-hypothetical, then each chunk is assigned to the bucket whose HyDE-vector it scored higher against. Because the retrieval pool is symmetric and the classification mirrors the retrieval direction, supporting/contradicting counts come out approximately balanced regardless of the actual distribution of evidence in the corpus (a topic that is 90% supported in the literature will still show a ~1:1 split here). Use fast mode for 'is there evidence on either side?', not for 'how is the field actually split?'. May also misclassify chunks that mention the topic but logically point the other way (e.g. a paper explaining 'BN is bad in transformers' may land in the contradicting bucket for an 'LN > BN' claim). 'deep': adds an independent per-chunk LLM NLI classification on top of the union pool, so counts reflect actual semantic distribution and can be arbitrarily asymmetric. ★ COST IS LINEAR IN POOL SIZE — one LLM call PER CHUNK, up to ~60 per invocation. Measured over 35 real calls: median 52 s end-to-end (the '~10s' this description used to claim was never the general case). That is money as well as time. Use 'deep' whenever classification accuracy or distribution shape matters — including controversy mapping and any analysis that interprets the supporting/contradicting ratio as a signal about the field."
  2. Changed1 schema field changed
    • removedInput schema / properties / detail / default
      Removed value: -"full"
  3. Changed1 schema field changed
    • changedInput schema / properties / detail / default
      Previous value: -"standard"New value: +"full"
  4. First observed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden — and it excels. It discloses the HyDE mechanism, the critical 'symmetric-by-construction' flaw in fast mode (counts are 'approximately balanced BY CONSTRUCTION and do NOT reflect actual literature distribution'), misclassification risks with a concrete example, and hard cost/latency data ('one LLM call PER CHUNK, up to ~60 per invocation,' 'median 52 s end-to-end'). This is unusually candid about failure modes and resource implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, method, classification output, use cases, then the two modes with the critical caveat flagged by 'IMPORTANT.' It is front-loaded with the core purpose before any mechanism detail, and the fast/deep contrast is structured so the caveat is unmissable. For a tool with this behavioral complexity, the length is justified rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 7-parameter tool with no output schema and no annotations, the description covers a remarkable amount: core semantics, HyDE mechanism, mode selection, the statistical caveat, cost/latency, and the selfContained return indicator. The main gaps are the two undocumented parameters (detail, categories) and the absence of a return-structure description beyond 'returns chunks with selfContained flag,' which a full output schema would have covered. These are meaningful but minor against the strong coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, above the low threshold, so the schema carries most of the parameter documentation burden. The description adds useful context for selfContainedOnly by explaining the selfContained flag as a 'safe-to-cite indicator' and enriches claim semantics. However, two parameters (detail and categories) are entirely undocumented in both schema and description, and the description adds no meaning for limit or run_id beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource pair ('Fact-check or substantiate a claim against the corpus') and immediately adds the distinguishing behavior: 'retrieves and CLASSIFIES evidence into supporting / contradicting / neutral groups.' This clearly separates it from sibling retrieval tools like search, find_related, and methodist_find, which do not classify evidence into stance groups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use context: 'Use for fact-verification, controversy mapping, 'is this claim known?' queries,' plus mode-level routing ('Use 'deep' when measuring controversy balance, literature distribution...' and 'Use fast mode for 'is there evidence on either side?', not for 'how is the field actually split?''). It stops short of naming sibling alternatives and stating when NOT to use this tool versus them, so exclusions are only implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.