Skip to main content
Glama
getsimba-ai

Simba MCP Server

Official
by getsimba-ai

Evaluate Study Run

evaluate_study_run

Assess a study run against a chosen policy, previewing the evaluation basis or submitting evidence, to see which checks pass or fail.

Instructions

Save an immutable assessment, or with preview=true see what one would contain without writing (nothing stored, no access event). The preview returns basis_hash, would_evaluate, would_not_evaluate (each with basis.reason and, when an earlier assessment of this run can lend the value, carry_forward_available; otherwise carry_forward_blocked with why: basis_changed, policy_changed, manual_signoff) and previous_assessment. Each submission is complete: a custom check has a value only if you supply it now in external_evidence or name it in carry_forward from an earlier assessment whose evidence basis and policy rules still match; a metric may not be in both. Submit finite numeric or strict boolean values with method and source_reference and the expected_basis_hash from the preview; stale evidence is 409 stale_evidence. External calculations are submitter-reported, not verified; carried rows keep carried_from provenance. Manual sign-off is never carried and requires a signed-in reviewer. Every unevaluated check carries basis.reason from a closed set and a suggested action; not_evaluated never passes. Choose policy_id explicitly; the report echoes policy_name and policy_was_newest. No automatic champion promotion. Built-in errors are fitted-window, not holdout; VAR remains unsupported.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
run_idYes
previewNo
policy_idYes
carry_forwardNo
external_evidenceNo
expected_basis_hashNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv0.7.1
    • addedInput schema / properties / carry_forward
      Added value: +{
      +  "anyOf": [
      +    {
      +      "additionalProperties": false,
      +      "description": "Reuse an earlier assessment's external evidence for the named custom metrics. Allowed only when the preview listed the metric under carry_forward_available for that assessment: same run, same evidence basis, same policy rules. Manual sign-off is never carried.",
      +      "properties": {
      +        "from_evaluation_id": {
      +          "type": "string"
      +        },
      +        "metrics": {
      +          "items": {
      +            "pattern": "^custom:[a-z][a-z0-9_]{0,63}$",
      +            "type": "string"
      +          },
      +          "maxItems": 20,
      +          "minItems": 1,
      +          "type": "array"
      +        }
      +      },
      +      "required": [
      +        "from_evaluation_id",
      +        "metrics"
      +      ],
      +      "type": "object"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Carry Forward"
      +}
    • changedInput schema / properties / external_evidence / anyOf
      Previous value: -[
      -  {
      -    "items": {
      -      "additionalProperties": false,
      -      "description": "Externally calculated numeric or strict boolean evidence; server applies policy. Manual sign-off is session-only and cannot be submitted with an API key. Source/method/digest are submitter-reported, not independently verified. Never submit a pass/fail status.",
      -      "properties": {
      -        "method": {
      -          "maxLength": 5000,
      -          "type": "string"
      -        },
      -        "metric": {
      -          "pattern": "^custom:[a-z][a-z0-9_]{0,63}$",
      -          "type": "string"
      -        },
      -        "source_reference": {
      -          "maxLength": 2000,
      -          "type": "string"
      -        },
      -        "source_sha256": {
      -          "pattern": "^[a-f0-9]{64}$",
      -          "type": "string"
      -        },
      -        "value": {
      -          "type": [
      -            "number",
      -            "boolean"
      -          ]
      -        }
      -      },
      -      "required": [
      -        "metric",
      -        "value",
      -        "method",
      -        "source_reference"
      -      ],
      -      "type": "object"
      -    },
      -    "type": "array"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]New value: +[
      +  {
      +    "items": {
      +      "additionalProperties": false,
      +      "description": "Externally calculated numeric or strict boolean evidence; server applies policy. Manual sign-off is session-only and cannot be submitted with an API key. Source/method/digest are submitter-reported, not independently verified. Never submit a pass/fail status. To reuse an earlier submission on the same basis, use carry_forward instead of retyping.",
      +      "properties": {
      +        "method": {
      +          "maxLength": 5000,
      +          "type": "string"
      +        },
      +        "metric": {
      +          "pattern": "^custom:[a-z][a-z0-9_]{0,63}$",
      +          "type": "string"
      +        },
      +        "source_reference": {
      +          "maxLength": 2000,
      +          "type": "string"
      +        },
      +        "source_sha256": {
      +          "pattern": "^[a-f0-9]{64}$",
      +          "type": "string"
      +        },
      +        "value": {
      +          "type": [
      +            "number",
      +            "boolean"
      +          ]
      +        }
      +      },
      +      "required": [
      +        "metric",
      +        "value",
      +        "method",
      +        "source_reference"
      +      ],
      +      "type": "object"
      +    },
      +    "type": "array"
      +  },
      +  {
      +    "type": "null"
      +  }
      +]
    • addedInput schema / properties / preview
      Added value: +{
      +  "default": false,
      +  "title": "Preview",
      +  "type": "boolean"
      +}
  2. Addedv0.5.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations. It discloses that preview writes nothing and creates no access event, that stale evidence returns 409 stale_evidence, that external calculations are submitter-reported and not verified, that carried rows keep carried_from provenance, that manual sign-off is never carried, and that built-in errors are fitted-window not holdout. These are behavioral traits an agent needs to know and are not present in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and information-rich, with the most important behavioral facts front-loaded (immutable save, preview mode, no write). Every sentence adds a distinct fact, but the density is high and some sentences are packed with multiple clauses that require careful parsing. It is appropriately sized for the tool's complexity, though slightly less structured than ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, preview vs. write modes, carry-forward rules, error conditions, provenance semantics), the description covers the critical operational context: what happens on write, what preview returns, what causes 409, what is never carried, and what is unsupported. The output schema exists, so return values need not be described in detail. The description is complete enough for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain the key parameter semantics: preview=true means no write, expected_basis_hash comes from the preview, carry_forward requires the metric to be listed under carry_forward_available, and external_evidence requires finite numeric or strict boolean values with method and source_reference. However, it doesn't explicitly map every parameter (e.g., run_id, policy_id) and doesn't explain the policy_id choice beyond 'Choose policy_id explicitly; the report echoes policy_name and policy_was_newest.' Still, the description adds substantial meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Save an immutable assessment, or with preview=true see what one would contain without writing.' This clearly distinguishes the two modes of the tool and names the resource (assessment of a study run). It also differentiates from siblings like list_study_evaluations and recommend_study_run by describing the write/preview behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: use preview=true to see what an assessment would contain without writing, and use the full submission to save. It also states exclusions: 'Manual sign-off is never carried and requires a signed-in reviewer,' 'VAR remains unsupported,' and 'No automatic champion promotion.' These exclusions help an agent decide when not to use this tool or what not to expect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools