Skip to main content
Glama
getsimba-ai

Simba MCP Server

Official
by getsimba-ai

Create Quality Policy

create_quality_policy

Define project-specific quality gates by creating an immutable policy with required checks, thresholds, and optional validation protocol, while preserving lineage via a source policy.

Instructions

Save project-specific checks. Policies are immutable: to change one, create a new policy, ideally with derived_from_policy_id naming the saved policy you started from (same study) so lineage is kept and the response carries diff (checks added/removed/changed by metric and field, protocol changes, name_changed, rationale_changed, rules_changed) — reordering checks is not a change. Built-in checks have metric (r_hat_max, mae, rmse, wape, prediction_mae, prediction_rmse, prediction_wape), maximum and required, and may instead use operator gte/between with minimum (default lte). Artifact-backed checks read what the fit already saved and need no external evidence: saved diagnostics (r_squared, mape, durbin_watson, pareto_k_pct, normality_p, loo_cv), retained-sampling counts (retained_chains, retained_draws_per_chain, ess_bulk_min, ess_tail_min, divergences) and provenance status (provenance:holdout, provenance:prior with expected review_required; blocked fails, absent is not_collected). Custom numeric checks use metric custom:, name, units, operator (lte/gte/between), applicable minimum/maximum and required. Boolean checks use kind=boolean, operator=equals and expected=true/false. Manual checks use kind=manual, equals, expected=true; agents can define these but cannot submit manual sign-off. Custom bounds may be negative. WAPE is a fraction. Prediction-window checks require saved finite actuals/predictions at unique dates after the saved training window; this does not certify untouched holdout provenance. No default thresholds are assumed. Declare at least one required check, use each metric once, and set maximum R-hat at least 1. At most 20 checks in total; the refusal names the count. A bad check is refused with one message naming its position and declared kind. Optional validation_protocol declares a temporal holdout split, configured sampling minima, R-hat and prediction WAPE limits before both runs launch under this policy. The backend validates policy rules.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameYes
checksYes
study_idYes
rationaleYes
validation_protocolNo
derived_from_policy_idNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changedv0.7.1
    • changedInput schema / properties / checks / items / description
      Previous value: -"Built-in metric + maximum, or custom numeric gate with metric custom:<slug>, name, units, operator and applicable bounds. Boolean/manual checks use kind, operator equals and strict boolean expected. Manual expected must be true and sign-off is session-only. At least one required gate per policy."New value: +"Built-in metric + maximum (or operator gte/between with minimum), or custom numeric gate with metric custom:<slug>, name, units, operator and applicable bounds. Boolean/manual checks use kind, operator equals and strict boolean expected. Manual expected must be true and sign-off is session-only. Rules over saved artifacts need no external evidence: saved diagnostics (r_squared, mape, durbin_watson, pareto_k_pct, normality_p, loo_cv; bounds may be negative for loo_cv and r_squared), retained-sampling counts (retained_chains, retained_draws_per_chain, ess_bulk_min, ess_tail_min, divergences; non-negative) and provenance status (provenance:holdout, provenance:prior; operator equals, expected review_required — blocked fails, an absent record is not_collected). Every threshold is the author's; an absent artifact never passes. At least one required gate per policy; at most 20 checks in total."
    • changedInput schema / properties / checks / items / properties / expected / type
      Previous value: -"boolean"New value: +[
      +  "boolean",
      +  "string"
      +]
    • changedInput schema / properties / checks / items / properties / metric / examples
      Previous value: -[
      -  "r_hat_max",
      -  "mae",
      -  "rmse",
      -  "wape",
      -  "prediction_mae",
      -  "prediction_rmse",
      -  "prediction_wape",
      -  "custom:benchmark_deviation"
      -]New value: +[
      +  "r_hat_max",
      +  "mae",
      +  "rmse",
      +  "wape",
      +  "prediction_mae",
      +  "prediction_rmse",
      +  "prediction_wape",
      +  "r_squared",
      +  "mape",
      +  "durbin_watson",
      +  "pareto_k_pct",
      +  "normality_p",
      +  "loo_cv",
      +  "retained_chains",
      +  "retained_draws_per_chain",
      +  "ess_bulk_min",
      +  "ess_tail_min",
      +  "divergences",
      +  "provenance:holdout",
      +  "provenance:prior",
      +  "custom:benchmark_deviation"
      +]
    • addedInput schema / properties / derived_from_policy_id
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Derived From Policy Id"
      +}
  2. Addedv0.5.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses far beyond sparse annotations: immutability of saved policies, refusal behavior ("the refusal names the count", "refused with one message naming its position and declared kind"), backend validation of rules, agent restrictions ("cannot submit manual sign-off"), provenance edge semantics ("blocked fails, absent is not_collected"), and response contents (diff carrying checks/protocol changes). No contradiction with readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose ("Save project-specific checks") followed by the most decision-relevant constraint (immutability). The description is long (~400 words) but each sentence carries distinct rules for a genuinely complex tool — check taxonomy, constraints, error behavior, protocol. Minor redundancy exists with the checks item description already embedded in the schema (e.g., "At least one required gate per policy; at most 20 checks"), so it is not maximally lean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with zero schema descriptions, the description is remarkably complete: every check kind and its expression, validation invariants (at least one required check, each metric once, max R-hat ≥ 1, ≤20 checks, no default thresholds), prediction-window semantics, validation protocol semantics, lineage/diff behavior, and error/refusal format. The output schema covers return values, so nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, putting the full burden on the description, and it delivers: categorizes the metric enum into built-in/artifact-backed/custom with per-kind parameter requirements (operator defaults like lte, custom:<slug> syntax, boolean vs manual shapes, expected semantics), explains validation_protocol as a temporal holdout split with configured sampling intent, and clarifies derived_from_policy_id lineage. This far exceeds the bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Save project-specific checks" states a specific verb (save/create) and resource (project-specific quality policy), and the immutability clause immediately establishes this as the creation tool in a create/list/get/diff/retire family. An agent can distinguish it from diff_quality_policies, list_quality_policies, get_quality_policy, and retire_quality_policy without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use context: "Policies are immutable: to change one, create a new policy, ideally with derived_from_policy_id naming the saved policy you started from (same study)" — this defines the derivation workflow clearly. However, it never names sibling alternatives (retire_quality_policy for retiring, diff_quality_policies for comparing) nor states explicit when-not-to-use conditions, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools