Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

did_summary

Read-only

Compare staggered difference-in-differences results across methods to assess robustness of treatment effects.

Instructions

One-call method-robustness comparison for staggered DID.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
yYesOutcome variable.
timeYesTime / period variable (integer-valued).
alphaNoSignificance level for confidence intervals.
groupYesUnit identifier.
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
clusterNoCluster variable for SE (defaults to ``group`` in each sub-method).
methodsNoMethods to run. Valid keys: ``'cs'``, ``'sa'``, ``'bjs'``, ``'etwfe'``, ``'stacked'``, or ``'all'`` / ``'auto'`` for all.auto
verboseNoPrint progress for each method.
controlsNoTime-varying covariates passed to methods that support them.
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathYesAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
first_treatYesFirst-treatment period per unit; NaN (or 0) for never-treated.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.
include_sensitivityNoIf ``True`` and ``'cs'`` is among the methods fit, compute the Rambachan-Roth (2023) *breakdown M\** -- the largest relative violation of parallel trends under which the treatment effect is still significantly different from zero. The value is added to ``model_info['breakdown_m']`` and to the ``breakdown_m`` column of ``detail`` (CS row only; other methods leave ``NaN``).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already signals this is a safe read operation. The description adds the behavioral fact that it performs a multi-method robustness comparison in one call, but provides no additional detail about internal execution, side effects, or requirements. The output schema covers return values, so this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused phrase with no filler or redundant content. It is front-loaded and every word contributes to the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The input and output schemas are rich, covering parameters and return structure, which reduces the burden on the description. However, the description leaves 'method-robustness comparison' undefined and does not help an agent choose among the many DID-related siblings, leaving meaningful gaps for a complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 16 parameters are documented in the input schema itself. The tool description adds no parameter-specific meaning, which matches the baseline of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific action (comparison) and resource (staggered DID methods), which goes beyond a tautology. However, it does not explicitly distinguish itself from sibling tools like did_summary_plot or did_report, so it lacks full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'one-call' implies a convenience use case for running multiple robustness checks in a single invocation, but there is no explicit statement of when to use this tool versus alternatives or any exclusions. Usage guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools