Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

cate_by_group

Read-only

Compute group-level average treatment effects from individual CATE estimates, grouping by a specified column or CATE quartiles, with confidence intervals and diagnostics for subgroup decisions.

Instructions

Group-level average treatment effects. Assumptions: Unconfoundedness given the covariates; Overlap / positivity across the covariate space; Nuisance functions are estimated consistently; cross-fitting controls overfitting bias. Pre-conditions: Covariates, a treatment indicator, and an outcome for each unit; Enough data to fit flexible nuisance models with sample-splitting / cross-fitting. Failure modes: CATE estimates are unstable or extrapolate beyond the covariate support -> Restrict to the overlap region, increase data, or use a doubly-robust learner (DR-/R-learner). Alternatives: sp.dml, sp.causal_forest, sp.tmle. Typical minimum N: 500.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
byYesColumn name to group by, or 'cate' to group by CATE quartiles.
alphaNoSignificance level for CIs.
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
resultYesResult from ``metalearner()`` / ``tarnet()``, a fitted ``causal_forest()`` model, or a raw array of per-unit CATE estimates.
n_groupsNoNumber of quantile groups when ``by='cate'`` or when the grouping variable is continuous.
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathYesAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description builds on that by explaining assumptions like unconfoundedness and overlap, as well as failure modes such as unstable estimates or extrapolation beyond covariate support. It also discloses that cross-fitting is used to control overfitting bias, which is useful behavioral context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, followed by assumptions, preconditions, failure modes, and alternatives. It is dense but efficient; a small deduction is warranted because some content, such as the specific alternative tool names, could be more actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The input schema and output schema already cover parameter semantics and return value shape, so the description does not need to repeat them. The description usefully adds assumptions, preconditions, failure modes, and a sample-size guideline. It is slightly incomplete in not giving clearer guidance for selecting among the many sibling CATE-related tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so the baseline of 3 applies. The main description does not add parameter-level detail beyond what the schema already documents, but the schema itself clearly explains 'by', 'result', 'detail', and the other parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Group-level average treatment effects,' which clearly identifies the resource and the operation domain. It is understandable on its own, but it lacks an explicit verb such as 'estimates' or 'computes,' and it does not explicitly distinguish this from sibling tools like cate_summary, cate_plot, or cate_group_plot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides assumptions, preconditions, failure modes, and a typical minimum N, which give useful context for when the tool is appropriate. However, it never explicitly states when to choose this tool over alternatives, and the Alternatives line lists names without conditions and with names ('sp.dml', etc.) that do not match the provided sibling tool list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools