Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

parallel_trends_plot

Read-only

Plot outcome means over time for treatment and control groups to visually assess parallel trends assumption before causal analysis.

Instructions

Plot raw outcome means over time for treatment and control groups.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
yYesOutcome variable.
axNoExisting axes to plot on.
ciNoShow 95% confidence intervals (+/-1.96 SE of mean).
idNoUnit identifier (for panel data).
aggNoAggregation function: 'mean' or 'median'.mean
timeYesTime period variable.
titleNoPlot title.
treatYesTreatment group indicator. Binary (0/1) for 2x2, or first-treatment-period for staggered (0 = never treated).
colorsNoColors for (treatment, control). Default: ('#E74C3C', '#2C3E50').
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
labelsNoCustom labels, e.g. ``{'treat': 'New Jersey', 'control': 'Pennsylvania'}``.
figsizeNoFigure size.
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathYesAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
treat_timeNoTreatment onset time. Draws a vertical line if provided.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, so the read-only nature is covered. The description adds that it plots raw (unadjusted) means, but it does not disclose behaviors like server-side result caching via as_handle or remote data fetching. There is no contradiction with annotations, but little additional behavioral context is added.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every word contributes to the core message. It is arguably too terse for an 18-parameter tool, but the brevity does not obscure the primary purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a rich 100%-covered schema and an output schema, the essential mechanics of calling the tool are present. However, the description does not connect this plot to its common use case (assessing parallel trends), nor does it distinguish it from the many sibling plot tools. Context such as staggered treatment handling is left entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 18 parameters in detail. The description only adds that y is aggregated as a mean, and it slightly oversimplifies by not acknowledging the 'median' option under agg. It provides no meaningful new parameter information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb and object: 'Plot raw outcome means over time for treatment and control groups.' It clearly states what the tool computes, and 'raw' helps separate it from model-based or robustness-oriented siblings. It does not explicitly name sibling tools, but its purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives such as parallel_trends_robustness, did_plot, or treatment_rollout_plot. The description states only what it does, leaving tool selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools