Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

honest_did

Read-only

Assess sensitivity of event-study estimates to parallel-trends violations, returning robust confidence intervals. Use when pre-trend tests reject, instead of abandoning the design.

Instructions

Rambachan-Roth (2023) 'honest' DID sensitivity analysis. Operates on a fitted event-study / DID result: pass result_id from a prior sp.event_study / sp.callaway_santanna / sp.did_imputation / sp.sun_abraham fit run with as_handle=true (or use the honest_did_from_result tool, which auto-extracts the event-study coefficients). Returns honest confidence intervals under varying degrees of parallel-trends violation (smoothness or relative-magnitude restrictions). Call this when a pre-trend test rejects at low power instead of abandoning the design. Validation: certified parity evidence.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
eNoEvent-time period (post-treatment horizon) to build the honest CI for.
alphaNoSignificance level for the honest CI.
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
m_gridNoGrid of M bounds on the parallel-trends violation to sweep.
methodNoRestriction family: 'smoothness' (second differences) or 'relative_magnitude'.smoothness
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathNoAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so no contradiction. The description adds that it operates on a fitted result and returns honest confidence intervals under parallel-trends violations, which is useful context. However, it doesn't clarify whether it requires a pre-existing result, what happens if no result_id is supplied, or the exact meaning of the 'Validation: certified parity evidence' statement, leaving some ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, reasonably front-loaded with the method and purpose, and the chaining instructions are compact. The final 'Validation: certified parity evidence' sentence is opaque and adds little value, slightly reducing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 parameters, output schema present, many DID siblings), the description covers the key context: what it does, when to use it, and how to get the needed fitted result. It does not explain how to interpret the outputs or guide parameter choices like m_grid, but the output schema and 100% parameter schema coverage relieve some of that burden. Overall adequate but with room for more practical guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds context for result_id (how to obtain it) and mentions smoothness vs relative-magnitude restrictions, which maps to the method parameter. It does not add detail for m_grid, alpha, e, or the common data-loading parameters, but the schema already documents those adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a Rambachan-Roth (2023) 'honest' DID sensitivity analysis that operates on a fitted event-study/DID result and returns honest confidence intervals. It names the specific prior-fit tools and explicitly contrasts itself with honest_did_from_result, so an agent can distinguish it from the many sibling DID-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Call this when a pre-trend test rejects at low power instead of abandoning the design.' It also tells the agent how to chain inputs (pass result_id from specific prior tools with as_handle=true) and points to honest_did_from_result as an alternative that auto-extracts coefficients. This is clear, though it doesn't enumerate negative cases or alternative sensitivity tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools