Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

general_bunching

Read-only

Detect behavioral responses to tax or benefit kinks by estimating excess mass in the density around a threshold, with bias correction and counterfactual testing.

Instructions

High-order bunching design with bias correction. Validation: validated evidence tier (known-truth, reference, external-parity, or Monte Carlo artifact). Assumptions: The counterfactual density would be smooth through the threshold absent the policy; Excess mass at the threshold reflects the behavioural elasticity of interest; No other discontinuity coincides with the threshold. Pre-conditions: A behavioural choice variable (earnings, hours, ...) with a known kink or notch in the budget/choice set; A visible empirical density of the running variable around the threshold. Failure modes: Round-number heaping or a coincident policy contaminates the bunching mass -> Exclude heaping points, widen the excluded region, and test the counterfactual polynomial order. Alternatives: sp.rdrobust, sp.rkd. Typical minimum N: 500.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNoRandom seed for reproducible stochastic steps.
alphaNoSignificance level for confidence intervals and tests.
cutoffNocutoff parameter (float).
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
n_bootNoNumber of bootstrap replications.
runningYesRunning variable (e.g. earnings).
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
bandwidthNoBandwidth used for local smoothing or kernel weighting.
bin_widthNoDefaults to bandwidth / 25.
data_pathYesAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.
polynomial_orderNoOrder of the counterfactual polynomial fit.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses validation tiers, explicit assumptions, failure modes (round-number heaping, coincident policy) and remediation steps (exclude heaping points, widen excluded region, test polynomial order). It also provides a typical minimum N of 500, which is actionable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured with clear labels (Validation, Assumptions, Pre-conditions, Failure modes, Alternatives, Typical minimum N). The first sentence carries the core purpose, and while the text is a wall of semicolon-separated clauses, no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not required. The description covers what the tool does, its assumptions, prerequisites, failure modes with remedies, alternatives, and sample-size guidance, making it essentially self-contained for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning by linking key parameters to the econometric setting: 'running' is a behavioural choice variable, 'cutoff' is the threshold of the kink/notch, and 'counterfactual polynomial order' appears in the failure-mode guidance. This is modest but real added value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific method ('High-order bunching design with bias correction') with a clear verb and object. The labels for Validation, Assumptions, Pre-conditions, Failure modes, and Alternatives distinguish it from generic bunching or RD tools, and it explicitly names alternatives (sp.rdrobust, sp.rkd).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Pre-conditions define the required setting: a behavioural choice variable with a known kink/notch and a visible density around the threshold. The 'Alternatives' line names sp.rdrobust and sp.rkd, but does not give explicit decision rules for when to choose them instead, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools