step-engineer
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@step-engineerOptimize sort in src/utils.py for speed, keep tests passing, and return a verified patch."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Step Engineer
A reusable, Step-5-Preview-specific agent harness for constrained search: explore better implementations and strategies inside a clearly defined environment, using measured feedback and fixed acceptance rules. Use its ready-made MCP server and CLI, or build your own interfaces around its reusable Python components. GPT, Grok, Claude, or another agent defines the task while Step explores the permitted tradeoffs.
The current implementation works on local implementation or policy files: it edits an isolated copy, measures candidates against fixed checks, preserves the best qualifying version, and returns a patch with verification evidence. A simulator can supply the environment and score; the deliverable is still reviewed code, not a deployed controller.
Your parent agent keeps its existing model and reasoning settings, including Ultra where supported. Step Engineer does not replace the parent agent, apply patches to the original project, or publish changes.
繁體中文說明 · Use cases · Task brief · Why Step? Evidence and charts · Build your MCP / CLI · Parent-agent instructions
When to delegate to Step
Use it when you can specify the environment, allowed decisions, objective, hard constraints, permitted tradeoffs, and final validation. A useful task has a feasible baseline and more than one plausible strategy. The parent fixes what counts as success; Step is free to search the implementation space.
Situation | Useful delegation | Readiness |
Fixed resource capacity | Select a higher-value combination of jobs under both CPU and memory limits | Runnable synthetic policy example; no live Step result claimed |
Behavior must stay identical | Reduce batch-processing or composition time while preserving exact outputs | |
Quality can trade against speed | Increase search throughput while keeping recall above a fixed threshold | Requires your own index workload, quality evaluator, and holdout data |
Competing scheduling objectives | Improve delivery profit or throughput under capacity and service constraints | Requires your own simulator and policy contract; no production controller is supplied |
Being near a limit is useful only when it improves the objective and remains valid on the intended workload. Real-world variation may require an explicit margin. Contradictory hard constraints require a revised task contract; Step must not silently relax them. See the use-case guide and copyable task brief before submitting a job.
Related MCP server: MCP ToolHub
Why a harness for Step?
Our specialization hypothesis is that Step is useful when explicit constraints, an objective evaluator, and room to explore better solutions meet. StepFun's GPU-kernel and training-data experiments motivate it. These are vendor evidence, not proof that Step outperforms other models or that proximity to a constraint is itself a measure of quality. We evaluate task outcomes, not rule exploitation.
Vendor-reported results, checked 2026-09-26: a fixed MLA workload on one H100, with a 24-hour budget per run. This compares best runs with different reasoning settings, not averages or equal-cost results. Official presentation; method, source data, and limitations.
Our own live case reduced fresh nine-pose composition time by 10.57%, with identical RGBA outputs in the tested cases. Two preceding attempts produced no patch. This is one local workload, not a comparative model benchmark. The evidence guide separates official results, our measurements, and untested hypotheses, and explains which model capabilities this harness currently uses.
How other agents use it
What the harness does, and why it exists
Step proposes the next change. The harness turns those proposals into a bounded, measured search: it controls the editable state, supplies execution feedback, and decides whether the saved result meets the caller's acceptance criteria. This runtime is the part developers reuse behind their MCP or CLI.
flowchart TD
O["Orchestrator defines environment, allowed tradeoffs<br/>Objective, hard constraints, and holdout checks"] -->|"MCP / CLI / tool adapter"| C
subgraph H["Step Engineer harness: reusable execution and validation"]
C["1. Validate scope and snapshot files<br/>Keep the original project intact"] --> B["2. Measure the unchanged baseline<br/>Establish the comparison"]
B --> L["3. Dispatch permitted tools<br/>Control reads, edits, and execution"]
L -->|"evaluate_candidate"| E["4. Run supplied checks and benchmarks<br/>Measure validity and improvement"]
E --> K["5. Save the best feasible candidate<br/>Keep measured progress"]
K -->|"Feedback for the next attempt"| L
L -->|"Normal loop stop"| V["6. Revalidate saved best; write evidence<br/>Accept a checked result, not the last draft"]
K -.->|"Saved snapshot"| V
G["Across the loop: time, tokens, tools, estimated cost<br/>Stop limits and final-validation time reserve"] -.-> L
end
L -->|"Context and tool results"| S["Step-5-Preview API"]
S -->|"Proposed edits and tool calls"| L
V --> R["Orchestrator reviews patch, metrics, usage, and stop reason"]The diagram shows the normal optimization path. An invalid baseline stops before model iteration. Cancellation or an execution error records a termination result; it does not guarantee final validation or an accepted patch.
Harness mechanism | Why it is needed | Implementation |
Validate the job, snapshot selected files, restrict edits | Give every trial a defined scope and preserve the original project | |
Measure a baseline and repeated candidate runs | Establish a comparable starting point; reduce reliance on a single favorable timing | |
Dispatch a fixed set of tools and sandbox supplied commands | Turn model requests into controlled execution with file, network, and output limits | |
Return measured tool feedback to Step | Let the next attempt respond to observed failures and scores | |
Save only a better feasible candidate | Preserve measured progress when later attempts regress or remain untested | |
Bound iteration and reserve time for final checks | Give exploratory work a stopping policy and leave room to verify the result | |
Recheck a fresh copy of saved best and write artifacts | Let the caller inspect the actual candidate, measurements, usage, and acceptance decision |
Ownership: the orchestrator supplies the objective, checks, benchmark, thresholds, and optional independent final checks. The harness executes these supplied evaluators and records runtime evidence; it does not design the tests or constitute a cross-model evaluation platform. Step proposes changes and reacts to feedback. The orchestrator reviews the result and decides whether to apply it.
Saving best is automatic after a qualifying measurement; restoring the working
candidate requires the restore_best tool. Final validation uses a fresh copy of
saved best, and accepted.patch contains changes only after acceptance. Without
separate final_checks, the result reports independently_checked=false. Estimated
cost limits and the validation-time reserve do not guarantee an exact bill or a
successful final check.
The final-validation allowance is deducted from both model-request and development command deadlines. Reaching that exploration deadline stops further tool dispatch and hands the saved best to final validation. Baseline and final checks use the remaining job deadline. The allowance retains a 40% job-time cap, and process cleanup adds overhead, so it remains a best-effort reservation.
Where MCP and CLI fit
MCP and CLI are entry points to this runtime. The built-in CLI calls Harness
directly; MCP and the documented custom wrappers use JobService for allowed source
roots, background jobs, polling, and cancellation. The caller owns the process/session
lifetime. Cloud models issue tool calls through a local host; they do not directly
access your filesystem. See custom interfaces.
Your starting point | Reuse this layer |
An agent with MCP support | Launch the built-in stdio server |
A terminal or automation script | Use |
Your own domain-specific MCP or CLI | Wrap |
An existing GPT, Grok, or Claude API tool loop | Use |
The current release provides reusable Python components and examples; it does not generate a new MCP/CLI project automatically. Candidate execution currently requires macOS. Larger-context and multimodal model capabilities are documented separately from what this text-based worker exposes.
What it provides
A Step API client with bounded requests, usage accounting, and no automatic retries.
An optimization loop with explicit file permissions, correctness checks, repeated benchmarks, and separate final validation.
A macOS process sandbox for fixed, parent-owned commands.
CLI and asynchronous MCP stdio interfaces.
Tool schemas and local dispatch for OpenAI Responses, Grok Chat Completions, and Claude tools.
A fully local, scripted demonstration that requires no API key.
The default worker setting is medium, with up to 16,384 output tokens and 300 seconds per request, and 250,000 total tokens per job. These are a starting configuration, not a guarantee of improvement. In one live workload, two high attempts exhausted their output limits without producing a patch; a medium attempt produced a verified improvement. See the case study for the failures, measurements, and limits of that comparison.
Requirements
macOS with
/usr/bin/sandbox-exec. Candidate execution fails closed on unsupported systems; there is no unsandboxed fallback.Python 3.11 or newer. Python 3.12 is recommended for the tested setup.
uv to install the locked dependencies.
For live optimization: a Step API key with model access and sufficient provider quota.
The sandbox supports the installed Python runtime and narrow system paths. Other runtimes or compilers may need additional setup. Validate the required toolchain before delegating a job; do not disable isolation to work around a missing runtime.
Start with the offline demo
Clone the repository, then run the demo:
git clone https://github.com/ShemYu/step-engineer.git
cd step-engineer
uv sync --frozen --python 3.12
uv run step-engineer run examples/batch_aggregation/job.json --offline-demoThis runs the real harness and sandbox with a handwritten, scripted solution. It demonstrates the pipeline; it is not a Step model benchmark. The example documentation explains the contract and verification.
Configure a live worker
Copy the credential template and restrict access:
cp .env.example .env.local
chmod 600 .env.localEdit .env.local locally and set the placeholder to your own key:
STEP_API_KEY=YOUR_STEP_API_KEY
STEP_BASE_URL=https://api.stepfun.ai/v1Never commit this file, paste the key into a model prompt, or include credential files in a job. The client accepts the documented Step endpoint only. STEPFUN_API_KEY is a fallback variable; an existing process environment takes precedence over --env-file.
uv run step-engineer --env-file .env.local doctor
uv run step-engineer --env-file .env.local run examples/batch_aggregation/job.jsondoctor checks local configuration and sandbox availability. It does not verify authentication, quota, or model access. Live submission sends selected source and tool feedback to StepFun and can incur charges.
The bundled live job uses medium, at most 12 model requests, 24 tool calls, 600 seconds, 250,000 total tokens, and an estimated US$0.50. Its per-request limits are 300 seconds and 16,384 output tokens.
Delegate from a parent agent
A local MCP host can launch the stdio server:
uv run --project /absolute/path/to/step-engineer step-engineer \
--env-file /absolute/path/to/private/step.env \
--runs-dir /absolute/path/to/optimization-runs \
serve --allow-root /absolute/path/to/your-projectAll paths above are placeholders. Repeat --allow-root for each authorized source root. If omitted, only the bundled example is allowed. No HTTP port is opened.
Tool | Input | Result |
|
| Starts a job and promptly returns its |
|
| Progress, termination state, and estimated cost |
|
| Measurements, validation, and artifact paths |
|
| Cancels an active local job |
Keep the MCP session alive for the job. Closing the server cancels its active jobs; restarting it does not resume them or send new model requests.
GPT, Grok, or Claude can be the parent, but a local application must execute their tool calls. A cloud model cannot directly access this machine's filesystem or localhost. The integration guide covers Codex, Claude Code, and the provider-neutral Python adapter. Adapter schemas and dispatch have local tests; paid end-to-end tests against all three parent providers are not claimed.
Define a job
Copy the example job, then inspect the full schema:
uv run step-engineer schemaField | Purpose |
| Source directory; relative to the job JSON in the CLI, absolute in MCP |
| Explicit UTF-8 files to copy; no directories, symlinks, hidden files, or traversal |
| Existing files the worker may change; a subset of |
| Non-editable files withheld from development tools until final validation |
| Required behavior, optimization target, and prohibited tradeoffs |
| Parent-owned correctness commands that must exit successfully |
| Parent-owned measurement command |
| Independent final validation; strongly recommended for real work |
| Metric name and |
| Numeric conditions required in every benchmark repetition |
| Repeated measurement count; scores use the median |
| Required final improvement, e.g. |
| Step worker effort: |
| Request, tool, token, time, estimated cost, and stagnation limits |
Commands use argv arrays, without shell expansion. {python} selects the installed interpreter; {workspace} and {scratch} expand to isolated directories. The command workspace is read-only. Write build products and temporary files to scratch.
elapsed_seconds is measured externally and includes process startup, input preparation, and cleanup. A trusted benchmark may emit custom metrics as its last JSON line:
{"metrics":{"recall":0.98,"qps":4200}}Program output cannot override external elapsed_seconds. Missing or non-finite metrics, failed checks, timeouts, and output/storage limit violations reject the measurement. A custom stdout score is still only as trustworthy as its evaluator.
Budgets and acceptance
Generic defaults are 12 model turns, 40 tool calls, 600 seconds, 250,000 total tokens, an estimated US$1, and four evaluations without improvement. Each request is capped at 16,384 output tokens and 300 seconds. The bundled example tightens the tool and estimated cost limits.
max_request_seconds accepts 5–600 seconds. Its effective limit is also capped by the job's remaining time after a final-validation reserve. Token and estimated cost reservations can stop a job before its nominal maxima. Longer output limits include reasoning tokens, not just visible answers.
Cost estimates use US$1 per million input tokens and US$2.70 per million output tokens, checked on 2026-09-26, and ignore cache discounts. They are conservative local estimates, not a billing guarantee. Interrupted or timed-out requests can still be billed; uncertain usage is marked, and requests are not automatically retried.
The harness first measures the unchanged snapshot. Only a qualifying, better evaluated candidate becomes best. Final validation uses a fresh copy of that saved version. An unmeasured last edit never replaces the saved best. Budget exhaustion and acceptance are reported separately: a run may exhaust its budget while an earlier measured candidate still passes final verification.
Improvement-threshold comparisons allow score-scale floating-point roundoff at the boundary, while still requiring a strictly better score. This tolerance does not relax metric constraints or compensate for benchmark noise.
max_no_improvement counts completed candidate evaluations that do not improve
the saved best, including infeasible candidates. A new best resets the count;
reads, writes, and rejected tool arguments do not increment it. Reaching the
limit stops the rest of the tool batch and proceeds to final validation.
The guard-repair case study records four bounded Step attempts and the subsequent developer-agent repairs and parent review, including failures and costs. It is not a model benchmark.
Artifacts are written under runs/<run_id>/ in the current working directory by default. Set --runs-dir explicitly when integrating a host:
Artifact | Meaning |
| Outcome, usage, measurements, and validation |
| Nonempty only when the saved best passes final acceptance |
| Diagnostic patch; not approved for application |
| Input hashes for checking source drift |
| Measurement and event records |
Inspect accepted, final_validation, and stop_reason together. Without separate final_checks, the harness reruns normal checks and the benchmark but reports independently_checked=false. The original source directory remains unchanged; applying a patch is a parent-owned decision.
Data and isolation boundaries
Live jobs send selected visible source contents and tool feedback to StepFun. Exclude credentials, unrelated data, and material you cannot share with that provider. Final-only files are withheld from the worker's development tools, but are visible to the program being tested during final validation.
The runner blocks network access, prevents writes to the command workspace, restricts filesystem reads, and constructs a clean child environment without API keys. Output is capped at 16 KiB per stream; scratch has a 16 MiB per-file limit and a monitored 64 MiB aggregate limit. Scratch is removed after commands.
This is a guardrail for personal engineering work, not a VM or a hostile multi-tenant execution service. Escaped descendants and tampering inside the same interpreter require stronger isolation. Source copies, reports, patches, or test output can contain sensitive project data; review them before sharing. Raw model reasoning is kept in memory for tool continuity, not saved in reports.
Development checks
uv sync --frozen --python 3.12
uv run pytest -q
uv run ruff check src tests tools examples/resource_allocationTests cover provider HTTP mocks, state transitions, path boundaries, real macOS sandbox execution, and MCP stdio. Passing them establishes harness behavior; it does not establish a model's quality on an unseen workload.
References
This project uses the MCP Python SDK v1 interface.
License
MIT. The private application assets described in the case study are not included or licensed by this repository.
Available Tools
4 toolscancel_optimizationA
Cancel a local active job and its current check. A provider request already sent may still be billed.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the mutation/non-destructive posture is already covered. The description adds genuinely new behavioral context beyond that: cancellation is local-only and cannot claw back an already-sent provider request, which may still be billed. That billing caveat is exactly the kind of consequence an agent needs and cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero padding, with the core action front-loaded and the billing caveat in the follow-up where it belongs. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter cancel tool with no output schema, the description covers the essential action and the key side effect (residual provider billing). It stops short of stating the resulting job state or idempotency behavior, but nothing critical to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required run_id parameter, and the description never mentions it or its expected format/origin. With low coverage the description is supposed to compensate, and it does not; the only mitigation is that run_id is self-describing from the sibling tools' naming.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Cancel a local active job and its current check.' This is unambiguously distinct from the sibling read/submit tools (submit_optimization, get_optimization_status, get_optimization_result), so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb – you cancel an active job rather than a completed one – but there is no explicit when-to-use, when-not-to-use, or alternative-tool routing. Adequate but leaves the agent to infer the preconditions (e.g., job must be active).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_optimization_resultARead-onlyIdempotent
Read measured results, final validation, and local patch/report paths. accepted=false means no changes are approved by the harness. Parent review is still required.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/closed-world, so the safety profile is covered. The description then adds genuinely useful semantics beyond them: 'accepted=false means no changes are approved by the harness' and 'Parent review is still required', which tells the agent how to interpret the payload.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, with the core read scope front-loaded and the interpretation caveat second. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description partially fills that gap by enumerating the return contents (results, validation, patch/report paths) and flagging the accepted-flag semantics. It is nearly sufficient, missing only the run_id provenance and any note on error/empty cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter run_id is never explained in the description – no hint about where the run_id comes from or what form it takes. With one required param and no compensating text, the agent must infer its meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and concrete resources: measured results, final validation, and local patch/report paths. It is clearly the results-retrieval tool, though it never names how it differs from the sibling get_optimization_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Timing is only implied – this reads the outcome of a run, so it belongs after a run completes – but there is no explicit when-to-use, no mention of the alternative get_optimization_status for in-progress checks, and no stated prerequisites for run_id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_optimization_statusBRead-onlyIdempotent
Read local job progress and estimated API cost. Poll at reasonable intervals.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered. The description adds legitimate extra context beyond that: it reports estimated API cost and advises polling cadence. It still omits whether polling has rate limits or whether the reported cost is a live or final estimate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the core capability is front-loaded ahead of the polling advice. It is efficiently sized for the tool's simplicity, though minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should convey more about return shape; it does say progress plus estimated cost at a high level. But it leaves 'local' undefined and does not explain the lifecycle relationship with get_optimization_result, which matters for this three-tool workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one required string parameter, run_id, and schema description coverage is 0%, so the description carries the burden of explaining it. It does not mention run_id at all, nor where the value comes from (presumably submit_optimization). The name is fairly self-evident, so this is only a modest gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('local job progress and estimated API cost'), which clearly separates it from submit_optimization and cancel_optimization. It does not explicitly distinguish itself from get_optimization_result, which is the closest sibling and the one an agent is most likely to confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Poll at reasonable intervals' implies the tool is meant for repeated status checks, which is useful implied guidance. However, it gives no explicit when-to-use vs. get_optimization_result, no when-not guidance, and no indication of what a terminal state looks like before switching to the result tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_optimizationA
Start a Step API optimization job in an isolated copy. Requires an allowed source_dir, explicit files, correctness checks, benchmark, and budget. Returns promptly with a run_id; poll status. Sends selected code to StepFun and may incur API charges. Never changes the parent model or original source.
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: the job runs in an isolated copy, returns promptly with a run_id, sends selected code to StepFun, may incur API charges, and never mutates the parent model or source. The egress/cost disclosure is exactly the kind of behavior an agent needs before invoking, and it is consistent with destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, no repetition, and front-loaded: what it starts, what it requires, what it returns, and the cost/safety implications.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a fire-and-forget submission tool with no output schema, the description adequately covers lifecycle (run_id + poll), cost, and non-destructiveness. The gap is that a very complex nested parameter object is left largely unexplained, which an agent configuring a benchmark job would need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% on a deeply nested JobSpec, so the description must carry the load. It names several required inputs (source_dir, files, checks, benchmark, budget), but leaves editable_files, direction, metric, repetitions, constraints, final_checks, and the entire Budget knob set undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Start) and resource (a Step API optimization job in an isolated copy), and implicitly distinguishes itself from the sibling status/result/cancel tools by describing a one-shot submission that returns a run_id to poll.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Prerequisites are enumerated clearly ('Requires an allowed source_dir, explicit files, correctness checks, benchmark, and budget'), and it routes the agent to polling after submission. It doesn't name get_optimization_status/get_optimization_result explicitly, so it stops short of full when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.1- First observed
cancel_optimization - First observed
get_optimization_result - First observed
get_optimization_status - First observed
submit_optimization
TDQS
Scored across 4 tools
Each tool maps to a distinct phase of the optimization job lifecycle: submit starts, get_status polls, get_result retrieves final output, cancel aborts. No two tools share the same action or resource, so an agent can easily choose the right one.
All four tools use snake_case with a consistent verb_noun pattern (submit_optimization, get_optimization_status, get_optimization_result, cancel_optimization). The structure is predictable and unambiguous.
Four tools cover the essential async job lifecycle without redundancy. This is a well-scoped set for a focused optimization service.
The surface covers the full lifecycle: create, poll status, fetch results, and cancel. No obvious operation is missing for managing an optimization job, and the run_id-centric design avoids dead ends.
Maintenance
Related MCP Connectors
Deterministic AI code review, with an audit record. Governance inside the agent loop.
Change-aware CI validation and affected-test guidance for coding agents.
Change-aware CI validation and affected-test guidance for coding agents.
Run verified read-only code tools: quant diagnostics + agent-ops preflight, no source exposure.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables deterministic evaluation of coding agents by exposing controlled repository tools and returning structured verification reports with pattern checks and repeat-run comparisons.MIT
- FlicenseAqualityAmaintenanceEnables coding agents to perform workspace-confined file operations, read-only Git inspection, and structured shell commands, while requiring out-of-band human approval for mutations and external executions and maintaining an audit trail.143-
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to audit generated code locally via LM Studio, using AST checks to catch syntax errors and dangerous calls, and enforcing APPROVED/REJECTED feedback loops for automatic refactoring.1MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to safely inspect, edit, and test code within a bounded repository environment to solve software engineering tasks and verify fixes.-