Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

benjamini_hochberg

Read-only

Apply Benjamini-Hochberg FDR correction to unadjusted p-values, controlling false discovery rate in multiple comparisons and producing certified parity evidence for decision-making.

Instructions

Benjamini-Hochberg (1995) FDR correction. Validation: certified parity evidence.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
pvaluesYesUnadjusted p-values.
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathNoAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the agent knows this is a read-only calculation. The description adds little behavioral context beyond naming the procedure; it doesn't disclose output shape, assumptions, or side effects. Since the annotation already covers the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences and contains no wasted words. The core identity is front-loaded; the validation note is second. It's appropriately minimal for a self-explaining statistical procedure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple FDR-correction function with only one required parameter (pvalues) and full schema coverage, the description plus schema are mostly sufficient. The main gap is that it doesn't relate to or distinguish the tool from the multiple-adjustment siblings (bonferroni, holm, romano_wolf), and it doesn't mention whether adjusted p-values are sorted or returned in original order. But output schema exists and the operation is a narrow read-only calculation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters are already documented in the schema. The description adds no parameter-specific meaning, but with full coverage the baseline of 3 applies. The detail parameter in the schema already explains payload depth options well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the Benjamini-Hochberg (1995) FDR correction, which is a specific verb+resource combination. It is distinguishable from siblings like romano_wolf, holm, and bonferroni because it names the exact procedure. However, it doesn't explicitly contrast with those FDR-adjustment siblings, so it loses the fifth point on sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states what the tool is but not explicitly when to use it compared to alternatives. The 'Validation: certified parity evidence' line hints at a use case (validated parity with an established implementation) but doesn't say where it fits among the many multiple-testing siblings. The detail parameter does explain when different payload depths are appropriate, providing a moderate usage signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools