Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

qdid

Read-only

Estimate quantile treatment effects by applying a difference-in-differences contrast to outcome quantiles in a binary group/time design, with bootstrap standard errors.

Instructions

Quantile Difference-in-Differences (QDiD): applies the DiD contrast to quantiles, [Q11(t)-Q10(t)] - [Q01(t)-Q00(t)], on a 2x2 design with bootstrap SE. This is NOT changes-in-changes -- Athey & Imbens (2006) propose CiC and explicitly criticise QDiD; use sp.cic for CiC. Validation: certified parity evidence. Assumptions: CIC rank invariance: the quantile rank in the untreated distribution is stable across groups; Continuous outcome support covering both groups in both periods; SUTVA (no cross-group spillovers). Pre-conditions: panel or repeated cross-section; group is binary 0/1; time is binary 0/1 (pre / post). Failure modes: Outcome heavily discrete / zero-inflated -> CIC rank-matching is unstable on discrete supports -- use QTE regression (sp.qte) or Firpo-RIF; Bootstrap CI across quantiles varies wildly -> Thin tails at extreme quantiles -- restrict to [0.2, 0.8] or raise n_boot to 2000. Alternatives: sp.qte, sp.did, sp.rifreg. Typical minimum N: 500.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
yYesOutcome
timeYesBinary pre / post indicator
alphaNoSignificance level for confidence intervals and tests.
groupYesBinary treated / control group
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
n_bootNoNumber of bootstrap replications.
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathYesAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
quantilesNoQuantiles to estimate, defaults to [0.1, ..., 0.9]
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description doesn't need to restate read-only. It goes well beyond by disclosing assumptions (CIC rank invariance, continuous support, SUTVA), pre-conditions, and failure modes. It also names the validation ('certified parity evidence') and gives typical minimum N. This is rich behavioral context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place. It front-loads the core definition and formula, then moves to critical differentiators (CiC), assumptions, pre-conditions, failure modes, and alternatives. No filler or repetition; it's structured as a decision aid. The length is justified by the method's complexity and the need to avoid misuse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage context, assumptions, pre-conditions, failure modes, alternatives, and sample-size guidance. It even mentions validation. Given that an output schema exists (context signals show it), the return format doesn't need to be described. The description is complete for an agent to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all 12 parameters are already documented. The description adds extra semantics by linking n_boot to a failure-mode remedy (raise to 2000) and quantiles to a restriction suggestion ([0.2, 0.8] for thin tails). It also explains the 'detail' parameter's purpose in agent planning. These additions go beyond the schema, so a 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise definition: 'Quantile Difference-in-Differences (QDiD): applies the DiD contrast to quantiles, [Q11(t)-Q10(t)] - [Q01(t)-Q00(t)]'. It names the method, gives the exact formula, and immediately distinguishes it from changes-in-changes, explicitly naming the sibling sp.cic. This is a specific verb-resource combination that an agent can unambiguously select.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit: it states the design (2x2), the pre-conditions (panel or repeated cross-section, binary group and time), and exactly when NOT to use it ('NOT changes-in-changes'). It provides specific alternatives (sp.qte, sp.did, sp.rifreg) and failure modes with concrete remedies, e.g., raising n_boot to 2000. An agent knows exactly when to call this tool vs. alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools