Skip to main content
Glama
getsimba-ai

Simba MCP Server

Official
by getsimba-ai

Assess Study Validation Pair

assess_study_validation_pair

Assess a validation pair to verify completed MMM runs, frozen inputs, sampling, R-hat, and prediction windows, returning blockers and an evidence hash for audit.

Instructions

Assess a validation pair and append a prediction-access audit event when evidence is available. Checks distinct completed MMM runs launched under the declared protocol, frozen inputs/settings/runtime, configured sampling, saved R-hat, declared prediction windows/WAPE and date coverage. Returns blockers and an evidence hash; does not fit, accept or promote. Includes saved retained chain/draw, ESS and divergence records when available, with null for older models. Optional prelaunch retained_sampling limits require complete native records and check chain/draw minima, bulk/tail ESS minima and maximum divergences; otherwise sampling_qualification is not_declared. Optional require_policy_review checks current signed-in analyst acceptance of each latest same-policy assessment, including freshness and rejection blockers. The holdout_provenance report distinguishes missing evidence, blocked version 1 full-input preprocessing and version 2 training-only preprocessing requiring further provenance review. The prior_provenance report checks which rows the recorded smart priors read and the frozen input hashes; priors without a record remain unavailable, and priors that read observations after the declared training end are blocked. fresh_validation provides replacement-window preflight for influence reports naming the full-model revision: later windows, replacement policy chronology, recorded prior exposure, retained diagnostics and provenance/review requirements. It never clears champion blocks. External business calculations and untouched holdout history remain unverified; decision_grade_ready stays false.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
study_idYes
policy_idYes
full_run_idYes
validation_run_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.5.0

TDQS

B3.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations declaring readOnlyHint=false, openWorldHint=true and idempotentHint=false, the description adds substantial behavioral context well beyond the annotations: it appends an audit event, returns blockers and an evidence hash, substitutes null for older models, gates optional checks (retained_sampling, require_policy_review) with fallbacks like 'not_declared', and explicitly states decision_grade_ready stays false. These disclosures are consistent with the non-read-only, non-idempotent annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is an unusually long, dense block of run-on sentences mixing many distinct topics (audit event, R-hat/WAPE, retained diagnostics, policy review, two provenance reports, fresh_validation). It is not front-loaded or scannable, and key routing/scoping information is buried mid-paragraph, so it exceeds what an agent needs to select and invoke the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be restated, and the description does cover behavioral outcome states and gating logic richly. However, it leaves parameters entirely unexplained and gives no explicit alternative-tool routing, so it is adequate but with real gaps for a 4-required-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and all four parameters (study_id, policy_id, full_run_id, validation_run_id) are undocumented in both schema and description. The prose references concepts like the 'declared protocol,' 'same-policy assessment,' and 'full-model revision' that gesture at the parameters, but it never defines what each id represents or how they interrelate, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Assess a validation pair and append a prediction-access audit event') and enumerates exactly what it inspects (MMM runs, frozen inputs, sampling, R-hat, WAPE, provenance). It also distinguishes itself from adjacent actions by stating it 'does not fit, accept or promote' and 'never clears champion blocks,' which separates it from launch/promote tools. It stops short of naming a sibling tool directly, so differentiation is inferred rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage conditions are implied heavily through the feature list (validation pairing, retainment, policy review, holdout/prior provenance, fresh_validation), but there is no explicit 'use this when / use X instead' routing to alternatives such as evaluate_study_run or get_study_validation_resolutions. The negative scoping ('does not fit, accept or promote') gives some when-not guidance, but an agent must infer the primary trigger from context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools