Skip to main content
Glama

Start a background simulation

start_job

Run long strategy evaluations, rankings, or stress scenarios in the background and get a job ID immediately. Use this to reach the 252-day horizon when direct calls cap at 60 days.

Instructions

Start a long run of evaluate_strategies, rank_strategies or run_stress_scenario in the background and get a job id back immediately. This is the ONLY way to run to the certified 252-day horizon; a direct call is capped at 60 days so it can answer inside a conversation. The arguments are checked before the job starts, and the response estimates its run time. At most 2 jobs run at once and the last 32 are kept, in this server's memory only. Poll with check_job.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
toolYesThe tool to run in the background.
argumentsNoThat tool's arguments, as a direct call takes them. days may go to 252. An unknown argument or a wrong type is refused before the job starts.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover only the safety profile (non-read-only, non-destructive, non-idempotent, closed-world). The description adds substantial behavior the annotations cannot convey: immediate job-id return, pre-flight argument validation, a run-time estimate in the response, a concurrency cap of 2 simultaneous jobs, retention of only the last 32, and that state lives in server memory only. This is exactly the added value the dimension rewards.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the primary action and the job-id return, then constraints, then the polling pointer. No filler and nothing repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but the description still flags the job id and run-time estimate. For a long-running, resource-limited, in-memory job launcher, the concurrency limit, retention count, and polling instruction cover everything an agent needs to call and follow up correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description still adds meaning: 'days may go to 252' clarifies the range limit that the schema does not encode, and it confirms validation semantics ('an unknown argument or a wrong type is refused before the job starts') for the pass-through arguments object.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (start) and resource (a background job running one of three named tools), and enumerates exactly which tools it wraps. An agent can distinguish it from evaluate_strategies/rank_strategies/run_stress_scenario directly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use rule ('the ONLY way to run to the certified 252-day horizon') and the reason the direct call is insufficient (capped at 60 days so it can answer inside a conversation). It also names the follow-up tool ('Poll with check_job'), leaving no routing inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.