Skip to main content
Glama

atlas_start_custom_eval_batch

Start batch evaluation of multiple candidates using a custom evaluation model (5 credits per candidate). Returns a batch_id. Poll with atlas_get_custom_eval_batch_status(batch_id) until status='completed', then fetch with atlas_get_custom_eval_batch_results(batch_id). Requires context_id from atlas_list_contexts, candidate_ids from atlas_list_candidates, and custom_model_id from the Atlas dashboard.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
context_idYesContext ID from atlas_create_context or atlas_list_contexts
template_idNo
detail_levelNostandard
candidate_idsYesCandidate IDs from atlas_list_candidates
custom_model_idYesCustom eval model ID (from Atlas dashboard)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds critical cost information (5 credits per candidate) not in annotations. Clarifies the async/non-blocking nature requiring polling. Annotations indicate readOnlyHint=false and idempotentHint=false; the description aligns by implying resource creation (returns new batch_id) without claiming idempotency or read-only status.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three well-structured sentences covering: 1) Purpose and cost, 2) Return value and polling workflow, 3) Prerequisites. No filler words; information density is high and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Excellent coverage of the complex async workflow and billing implications for a batch operation. Minor gap: does not explain optional parameters (template_id, detail_level) or what constitutes brief vs standard vs deep evaluation levels, though the async workflow explanation compensates for lack of output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 60% schema coverage, the baseline is 3. The description mentions the 3 required parameters but largely repeats schema descriptions (context_id from atlas_list_contexts, etc.) without adding usage context. It completely omits the 2 optional parameters (template_id, detail_level), failing to compensate for schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific action (Start batch evaluation), resource (multiple candidates using custom evaluation model), cost (5 credits per candidate), and distinguishes from siblings by specifying the custom model workflow vs other batch operations like atlas_start_batch_gem.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly describes the async lifecycle: returns batch_id, poll with atlas_get_custom_eval_batch_status, fetch results with atlas_get_custom_eval_batch_results. Also lists prerequisites requiring context_id, candidate_ids, and custom_model_id from specific sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources