Skip to main content
Glama

atlas_set_custom_eval_rubric_overrides

Idempotent

Set rubric overrides for a custom eval model. Overrides let you customize the AI-inferred rubric (adjust weights, rename dimensions, add scoring criteria) without re-inferring. Replaces any existing overrides. model_id from atlas_create_custom_eval_model or atlas_list_custom_eval_models. Free.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
model_idYesModel ID from atlas_create_custom_eval_model or atlas_list_custom_eval_models
overridesYesOverride object — keys are dimension names, values are override configs (weights, criteria, etc.)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare idempotentHint=true and destructiveHint=false; the description adds valuable context by clarifying the replacement semantics ('Replaces any existing overrides') and disclosing cost ('Free'). No contradictions with annotations. Could mention validation errors or requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five sentences efficiently structured: action definition, capability explanation, behavioral warning, parameter sourcing, and cost. Every sentence earns its place with no redundancy or filler. Front-loaded with the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a setter tool with 2 parameters and no output schema, the description adequately covers the operation's purpose, side effects (replacement), prerequisites (model_id source), and cost. Lacks explicit error handling or validation details, but sufficient for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, establishing baseline 3. The description adds conceptual context for the overrides parameter (explaining it customizes 'AI-inferred rubric' with specific operations like adjusting weights/renaming dimensions) and reinforces model_id sourcing. Elevates slightly above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'Set rubric overrides for a custom eval model' with specific capabilities (adjust weights, rename dimensions, add scoring criteria). It distinguishes from siblings by emphasizing 'without re-inferring' (contrasting with atlas_infer_custom_eval_rubric) and 'replaces existing overrides' (contrasting with clearing or appending).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context on when to use (customizing without re-inferring) and warns about replacement behavior ('Replaces any existing overrides'). Mentions valid sources for model_id. Could be improved by explicitly stating when to use atlas_clear_custom_eval_rubric_overrides instead, but the contrast is implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources