Skip to main content
Glama

Optimize Submit

optimize_submit

Optimize continuous design variables to meet a performance objective and constraints, then report whether the final point is proven within measurement tolerance.

Instructions

Vary parameters until the spec is met, then say whether it was PROVEN — the step that closes the design-to-spec loop.

Everything else in this family measures; this searches. Two things make it different from a generic minimizer, and both come from the performance-contract layer underneath it:

A constraint verdict has three states. indeterminate — the measurement's uncertainty band straddles the limit — is NOT a failed step. An optimizer that reads it as a failure walks away from good designs; one that reads it as a pass converges on unproven ones. During the search an indeterminate constraint is scored on its nominal value, so it neither attracts nor repels, and the winner is proved properly at the end.

Convergence is not proof. A simplex can settle on a point that clears its limit by 2 % while its own grid-convergence band is 5 % wide — noise with a favourable sign. proven is therefore reported separately from converged, and is True only when the final measurement has every constraint at pass and no margin swallowed by its own band.

variables must be continuous and bounded — {"name": "diameter_mm", "min": 5, "max": 25, "start": 10}. An optimizer without a box walks to values that satisfy the arithmetic and mean nothing physically.

objective and each constraints entry use the performance-requirement mapping ({tool, metric, conditions, limit, screen, band_pct}), with "$<variable>" in conditions carrying the candidate's value. tier='auto' searches cheaply on each block's screen estimator, then polishes on the real tool from where the screen landed; 'screen' or 'solver' runs just that leg.

The search is a bounded Nelder-Mead — derivative-free because there is no adjoint through a CFD solve — so every evaluation is a real measurement. budget ({max_evals, max_wall_s}, default 40 evaluations) is the ceiling; revisited points are served from cache and do NOT count against it.

Shape optimization. Pass a recipe (+ fixed_inputs) and it is rebuilt for every candidate, so the search varies GEOMETRY rather than only numbers — "$handle" in a response's conditions is that candidate's part, and the returned history carries the handle each point built. Pass handle instead to optimize parameters against one fixed part; pass neither and the search is purely parametric.

Geometry-driven candidates are built on the MAIN thread through the worker's work queue, because FreeCAD's document API is not thread-safe and this search runs as a background job. That queue is drained once per incoming request, so a shape search only advances while you are polling job_status/job_result — the poll you must do anyway is what gives it its turn. Poll at your normal cadence and it simply works; stop polling and it stalls rather than finishing in the background. A candidate whose recipe fails to build is scored out as an infeasible point, not an error. study_submit remains the right tool for a FIXED grid over recipe geometry, which needs no queue at all.

Returns {job_id, status}; poll job_result for {ok, proven, stop_reason, best_params, best_value, objective: {name, metric, sense, value, band_pct}, constraints: [{name, state, measured, limit, band_pct, margin, margin_pct, detail, trust_reasons?}], phases: [{tier, n_evals, best_params, best_value, converged, reason}], history: [{i, tier, params, value, score, feasible, cached}], n_evals, n_cached, budget, variables, warnings}.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
tierNoauto
budgetNo
handleNo
recipeNo
objectiveYes
variablesYes
constraintsNo
fixed_inputsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the minimal annotations (readOnlyHint false, destructiveHint false) by disclosing nuanced behavior: three-state constraint verdicts, convergence not equaling proof, bounded Nelder-Mead with derivative-free search, budget ceilings, cached evaluations not counting, shape candidates building on the main thread, and the search stalling unless polled. This is exceptional transparency for a complex background-job tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but its length is justified by the tool's complexity and the complete absence of schema descriptions. It is well-organized with bold headers and front-loads the core purpose and key differentiators. A few sentences are dense enough that they could be tightened, but no significant part is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters, nested objects, no output schema, and a complex asynchronous optimization workflow, the description is remarkably complete. It specifies the returned job_id/status, the full job_result structure with phases, history, constraints, trust_reasons, n_evals, n_cached, budget, variables, and warnings, so an agent has everything needed to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the full burden, and it does. It explains variables with a concrete example, defines objective/constraints via the performance-requirement mapping, clarifies tier ('auto', 'screen', 'solver'), budget semantics with defaults, and the meaning of recipe, fixed_inputs, and handle. Every parameter is given meaningful context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a crisp, specific statement: 'Vary parameters until the spec is met, then say whether it was PROVEN — the step that closes the design-to-spec loop.' It clearly identifies this as an optimization submission tool and explicitly contrasts it with the rest of the family ('Everything else in this family measures; this searches'), making it unmistakable from siblings like study_submit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit and thorough. It says when to use this tool versus siblings ('Everything else in this family measures; this searches'), and specifically names study_submit as the right tool for a fixed grid over recipe geometry. It also gives detailed selection rules among tier values, parametric vs shape optimization, and handle vs recipe vs neither.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools