io.github.TheTsungYing/annealbridge
This server solves structured combinatorial optimization problems (binary/bounded-integer variables, linear/quadratic objectives, hard/soft linear constraints) and returns ranked, re-validated solutions.
Solve problems (knapsack, assignment, scheduling, TSP, subset selection) via solve_optimization, choosing among eight local/remote backends (exact, simulated_annealing, tabu, simulated_bifurcation, D-Wave QPU/hybrid, Fujitsu DA).
Validate a problem before solving (validate_optimization_problem) to catch semantic errors, warnings, and estimate compiled size.
Get server capabilities and the full problem JSON schema (get_optimization_capabilities), including backend availability, enabled status, limits, and supported variable types/operators.
Get advisory backend rankings for a problem without solving (recommend_backend), with reason codes and blocking errors.
Receive structured outcomes: success, infeasible (with diagnostics), invalid_problem, backend_unavailable, resource_limit_exceeded, etc., never exceptions.
Inspect returned solutions with per-constraint evaluations, objective values, soft violation scores, and optimality/infeasibility proofs when using the exact backend.
Provides integration with the Fujitsu Digital Annealer as a remote backend for solving combinatorial optimization problems via its QUBO API V4 over HTTPS.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@io.github.TheTsungYing/annealbridgeI can carry 10kg. A(10,6) B(8,5) C(7,4) D(6,3). Which items maximize value?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AnnealBridge
English | 繁體中文
Combinatorial optimization middleware for AI agents. An agent describes what to optimize as structured JSON; AnnealBridge decides how to encode and solve it, checks every answer against the original problem, and returns ranked, verified solutions over MCP, a CLI, or plain Python.
flowchart LR
U[Natural language] --> A[AI agent]
A -->|OptimizationProblem JSON<br/>variables · objective · constraints| B
subgraph B[AnnealBridge]
direction LR
V[validate] --> C[compile<br/>BQM / CQM] --> S[solve<br/>local or remote] --> R[re-validate against<br/>the original problem] --> K[rank]
end
B -->|SolveResult<br/>ranked, verified solutions| A
A --> N[Natural-language answer]Quick start
pip install annealbridgeA 0/1 knapsack: four items, capacity 10, maximize value. No file needed.
from annealbridge.models import OptimizationProblem
from annealbridge.orchestration import OptimizationService
problem = OptimizationProblem.model_validate({
"version": "1.0",
"name": "knapsack",
"variables": [{"name": n, "type": "binary"} for n in ["a", "b", "c", "d"]],
"objective": {"direction": "maximize", "linear_terms": [
{"variable": "a", "coefficient": 10}, {"variable": "b", "coefficient": 8},
{"variable": "c", "coefficient": 7}, {"variable": "d", "coefficient": 6}]},
"constraints": [{"id": "capacity", "type": "hard", "operator": "<=", "rhs": 10, "terms": [
{"variable": "a", "coefficient": 6}, {"variable": "b", "coefficient": 5},
{"variable": "c", "coefficient": 4}, {"variable": "d", "coefficient": 3}]}],
})
result = OptimizationService().solve(problem)
print(result.status) # success
print(result.solutions[0].variables) # {'a': 1, 'b': 0, 'c': 1, 'd': 0}
print(result.solutions[0].objective_value) # 17.0Domain failures come back as results, never as exceptions: result.status
is one of success, infeasible, invalid_problem,
resource_limit_exceeded, backend_unavailable, configuration_error or
solver_error. See docs/output-format.md.
Optional extras:
pip install "annealbridge[mcp]" # + MCP server (annealbridge-mcp)
pip install "annealbridge[dwave]" # + D-Wave cloud backends
pip install "annealbridge[all]" # everything
pip install "annealbridge[gpu]" # + PyTorch, for simulated_bifurcation on CUDA[gpu] is deliberately not part of [all]: PyTorch is a large download,
and on Windows the wheel PyPI serves is the CPU-only build, so a CUDA run needs
torch installed from PyTorch's own index first — see
docs/backends.md.
Related MCP server: Gurddy MCP Server
Use it from an AI agent (MCP)
With uv installed, add the server to
claude_desktop_config.json (or your host's equivalent) and restart the
host; where Claude Desktop, Claude Code and Codex keep that configuration is
listed in docs/mcp.md.
The first run fetches the package into uvx's own cached environment; that
download — numpy, dimod, dwave-samplers and the rest — can take tens of
seconds, long enough for a host's start-up timeout to show the server as
disconnected. Warm the cache once in a terminal first:
uvx --from "annealbridge[mcp]" annealbridge-mcp --versionIt resolves the environment and prints the version; from then on the host starts from that cache.
{
"mcpServers": {
"annealbridge": {
"command": "uvx",
"args": ["--from", "annealbridge[mcp]", "annealbridge-mcp"]
}
}
}Claude Code registers it in one line:
claude mcp add annealbridge -- uvx --from "annealbridge[mcp]" annealbridge-mcpOptimizing is then an ordinary chat:
You: I can carry 10 kg. Item A is worth 10 and weighs 6, B is worth 8 and weighs 5, C is worth 7 and weighs 4, D is worth 6 and weighs 3. Which ones should I take?
Behind the reply, the agent writes the request as problem JSON and calls
solve_optimization. Every solution comes back re-validated against the
original constraints, with optimality_proven: true on the exhaustive
exact backend; a document the server rejects comes back as
invalid_problem with every error and a fix for each, which the agent
applies before sending it again. Before solving on a remote backend, or on
a large problem, it calls validate_optimization_problem first, so a mistake
costs nothing; it calls get_optimization_capabilities when it needs the
backend list or the full schema.
Agent: Take A and C: value 17 at exactly 10 kg. The runners-up are A and D (16, at 9 kg) and B and C (15, at 9 kg). This is the proven optimum.
The wording is the agent's; the numbers are the tool result. Another tool,
recommend_backend, ranks the backends for a problem and is advisory only;
when you did not name a backend and more than one local backend fits, the
server instructions tell the agent to show the top entries and ask which to
run rather than to decide for you.
If your host lists prompts in its input menu, three of them — pick_subset,
assign and schedule_shifts — walk the agent from your own sentence to a
problem document of that everyday shape, which it then solves. The server
also serves resources: the five
example documents under annealbridge://examples/ and the full problem
schema at annealbridge://schema, so an agent can read a complete example or
the schema itself instead of guessing.
What it is not for. AnnealBridge does not handle continuous
(real-valued) variables, non-linear objectives or non-linear constraints.
Only exact proves that an answer is optimal or that no feasible one
exists, and by default it takes at most 24 compiled variables (slack bits
included); the other local backends (simulated_annealing, tabu,
simulated_bifurcation) are heuristics whose best answer may not be the
optimum. See the
Backends table below and
docs/limitations.md.
Any stdio-capable MCP host works the same way, a streamable-http transport
exists, and pipx or a pip-installed server behind an absolute path work in
place of uvx. uvx reuses the environment it resolved on its first run, so
a new release reaches an existing install only after uv cache clean annealbridge and a host restart; see docs/mcp.md.
Use it from the command line
Save the problem JSON below as knapsack.json, then:
annealbridge solve knapsack.jsonProblem: knapsack
Backend: exact
Status: success
Attempts: 1
Elapsed: 2.7 ms
Best solution (rank 1)
objective (maximize): 17
soft violation score: 0
item_a = 1
item_b = 0
item_c = 1
item_d = 0
Hard constraints: 1 / 1 satisfied
Soft constraints: 0 violations
Optimality proven: yesElapsed is the service's own wall clock and varies from run to run. Add
--json for the full SolveResult, --backend simulated_annealing to
override the backend, or try validate, recommend, capabilities,
example and export-schema. See docs/cli.md.
The problem JSON
The document behind the MCP and CLI examples above, the reduced form of examples/knapsack.json:
{
"version": "1.0",
"name": "knapsack",
"variables": [
{"name": "item_a", "type": "binary"},
{"name": "item_b", "type": "binary"},
{"name": "item_c", "type": "binary"},
{"name": "item_d", "type": "binary"}
],
"objective": {
"direction": "maximize",
"linear_terms": [
{"variable": "item_a", "coefficient": 10},
{"variable": "item_b", "coefficient": 8},
{"variable": "item_c", "coefficient": 7},
{"variable": "item_d", "coefficient": 6}
]
},
"constraints": [
{
"id": "capacity",
"type": "hard",
"terms": [
{"variable": "item_a", "coefficient": 6},
{"variable": "item_b", "coefficient": 5},
{"variable": "item_c", "coefficient": 4},
{"variable": "item_d", "coefficient": 3}
],
"operator": "<=",
"rhs": 10
}
],
"solver": {"backend": "exact"}
}Integer variables ("type": "integer" with bounds, "version": "1.1"),
quadratic objective terms, soft constraints with weights and per-backend
solver preferences are described in
docs/problem-format.md.
annealbridge export-schema prints the JSON Schema an agent can use for
structured output.
Five ready-to-run examples live in the repository —
knapsack,
assignment,
TSP,
integer knapsack and
shift scheduling.
The installed package carries the same files: annealbridge example lists
them and annealbridge example knapsack > knapsack.json saves one (or pipe it
straight in with annealbridge example knapsack | annealbridge solve -); the
MCP server serves the same five documents as the resources under
annealbridge://examples/.
How it works
The agent produces an OptimizationProblem: binary or bounded-integer
variables, a linear or quadratic objective, and hard or soft linear
constraints. Nothing else. AnnealBridge then, deterministically:
validates the problem and collects every error in one pass;
compiles it into a BQM or a CQM, computing penalties, slack and integer encodings itself;
solves it on a local or remote backend;
re-validates every candidate against the original JSON, never trusting solver energy;
ranks the feasible solutions and returns the top K with per-constraint evaluations.
An opt-in solver.postprocess step can also repair and locally improve the
best samples in the original variables before ranking; what it produces is
re-validated like any other candidate and marked by each solution's source.
An optional solver.wall_clock_limit_seconds caps how long a solve on a local
heuristic backend may take: when it runs out, the result holds what was
completed, still re-validated and ranked, and says so
(wall_clock_limit_reached). An MCP client that cancels a solve stops it
rather than leaving it to run on.
The agent never writes a QUBO matrix, a penalty weight, a slack variable or an integer encoding, and every step is testable without an AI, a network or a vendor account.
Backends
Eight backends sit behind one protocol.
Backend | Kind | Path | Notes |
| local | BQM | Enumerates every assignment; 24 compiled variables by default |
| local | BQM | Heuristic; the general-purpose choice, best of the three on small or hard-constrained problems; honours |
| local | BQM | Heuristic multistart tabu search, strong on dense QUBOs; the first choice once a dense problem is large; honours |
| local | BQM | Heuristic dense-matrix dynamics (Goto et al. 2021), the choice for large dense unconstrained QUBOs, weaker on problems whose hard constraints compile to penalties; honours |
| remote | BQM | D-Wave quantum annealer via |
| remote | BQM | D-Wave Leap hybrid BQM solver |
| remote | CQM | D-Wave Leap hybrid CQM solver; native constraints |
| remote | BQM | Fujitsu Digital Annealer, QUBO API V4 over HTTPS, no SDK |
Remote backends need their vendor credential and
ANNEALBRIDGE_ALLOW_REMOTE=true; without both they report
backend_unavailable. annealbridge recommend ranks the backends for a
problem without solving it and never changes the one you asked for. Setup
and per-backend behaviour: docs/backends.md.
Design guarantees
Business-level contract in both directions. Variables, objective and constraints in; ranked solutions with per-constraint evaluations out. No solver internals leak either way.
Two compiler paths. BQM (automatic penalties, binary slack, encoded integers) for annealers; CQM (native constraints and integers) for the Leap hybrid CQM solver. The backend chooses by declaring what it supports.
Bounded integers without exposure.
"version": "1.1"adds integer variables with the encoding hidden;1.0behaviour is pinned by a golden test.Structured failures, never exceptions. Every outcome is a
SolveResultwith astatus; every failure carries a stable error code with arecommended_action.infeasibleis an answer, not a failure.No silent decisions. An unavailable backend is reported, never swapped. An over-limit parameter is rejected, never clamped. An undeclared field is rejected, never ignored. Validator warnings travel with every solve result.
Safe by default. Remote execution and remote retries are off until enabled; every limit is an environment variable enforced as an error; vendor credentials are redacted from results, logs and error messages. The streamable-http transport has no authentication; keep it on a private network. See docs/security.md and SECURITY.md.
Enforced architecture. Import boundaries, "no backend names in the orchestration, validation or interface layers", and "a new backend plugs in without touching the pipeline" are tests, not conventions.
Documentation
The pages below live under docs/.
Page | What it covers |
The input JSON: variables, objective, constraints, solver preferences | |
| |
Error catalog, warning codes, reason codes, exit codes | |
The | |
The MCP server: tools, prompts, resources, host configuration, Inspector | |
The eight backends, D-Wave and Fujitsu setup, adding a backend | |
Every | |
Layers, package layout, design principles | |
Defaults, limits, credential redaction, what reaches a vendor | |
Test layout, golden tests, live tests, CI | |
Known limits and what is out of scope |
Development
git clone https://github.com/TheTsungYing/AnnealBridge.git
cd AnnealBridge
pip install -e ".[all,dev]"
pytestpytest runs the full suite with no skip and no xfail and never touches the
network; the live vendor tests are opt-in (pytest -m remote). To install the
development version without a checkout:
pip install "annealbridge[all] @ git+https://github.com/TheTsungYing/AnnealBridge.git".
Architecture rules, design principles and the pull-request checklist are in
CONTRIBUTING.md.
Version 0.3.0: the problem contract (1.0 / 1.1), the eight backends, the
CLI and the MCP tools are complete and covered by tests. What is not
supported, by design for now, is listed in
docs/limitations.md;
changes are in CHANGELOG.md.
License
Available Tools
4 toolsget_optimization_capabilitiesA
Describe what this optimization server accepts and which solver backends are usable right now.
Call this when you need the supported problem vocabulary (variable types,
constraint operators, objective terms, accepted schema versions) or, per
backend, whether it is installed/configured (available), whether server
policy permits it (enabled), and its resource limits. It is not required
before every problem: for a small binary problem on a local backend the
example in the server instructions already shows the document shape. This
performs no solving and no network requests.
The full problem JSON schema is several times the size of everything
else, so problem_json_schema is null unless include_schema is true. Ask
for it when the document needs more than the example shows — integer
variables, soft constraints, solver preferences — or before inventing a
field; every field description is in it.
supported_variable_types lists the variable types a problem may declare —
"binary" and "integer" — and schema_versions lists every problem schema
version this server accepts, newest last. Integer variables are only
allowed when the problem carries "version": "1.1" at its top level;
"version": "1.0" accepts binary variables only.
| Name | Required | Description | Default |
|---|---|---|---|
| include_schema | No | Whether to include the full OptimizationProblem JSON Schema as problem_json_schema. The default false keeps the response small; pass true only when the schema itself is needed, e.g. for integer variables or an unfamiliar field. |
Output Schema
| Name | Required | Description |
|---|---|---|
| backends | Yes | Every registered backend, with its flags and limits. |
| schema_version | Yes | The newest problem schema version this server accepts; use it unless an older one is needed. |
| schema_versions | Yes | Every accepted problem schema version. Derived from the model, not hard-coded; "1.1" is a superset of "1.0". |
| problem_json_schema | Yes | The full OptimizationProblem JSON Schema, identical to what annealbridge export-schema prints, when it was requested (include_schema on the MCP tool); null otherwise. Its field descriptions explain the document one level down. |
| annealbridge_version | No | The installed package version that produced this view; "unknown" outside an installed distribution. |
| supported_variable_types | Yes | The variable types a problem may declare. |
| supported_objective_terms | Yes | The kinds of objective term a problem may carry. |
| supported_constraint_operators | Yes | The operators a constraint may use. |
| inequality_requires_integer_coefficients | Yes | Whether <= / >= constraints need integral coefficients and rhs, which the slack encoding requires. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to rely on, the description carries the full burden and does so thoroughly. It discloses that the tool 'performs no solving and no network requests,' explains that problem_json_schema is null unless include_schema is true, and even surfaces version-specific behavior such as integer variables only being allowed with version 1.1.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Each of the three paragraphs earns its place: purpose and usage, optional-parameter rationale, and version-specific constraints. The content is front-loaded with the core purpose, and no sentence is filler or redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple parameter surface and the presence of an output schema, the description covers everything needed: when to call it, what it returns, when to opt into the larger schema, and behavioral guarantees. It also differentiates the tool from the sibling tools without needing to enumerate their schemas.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the single include_schema parameter well, and the description adds further decision-useful meaning: when to pass true (integer variables, soft constraints, solver preferences, unfamiliar fields) and why false is the default (keeping the response small). This goes beyond the schema's own description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Describe what this optimization server accepts and which solver backends are usable right now.' It also distinguishes itself from the solving-focused sibling tools by stating 'This performs no solving and no network requests,' so an agent can identify it as a read-only capabilities query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to call the tool ('Call this when you need the supported problem vocabulary...') and when it is unnecessary ('It is not required before every problem'). It also gives precise guidance for the optional include_schema parameter, telling the agent to request the full schema when handling integer variables, soft constraints, solver preferences, or unfamiliar fields.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_backendA
Rank the solver backends of this server for a given problem, without solving it.
Advisory only: solve_optimization always uses problem.solver.backend exactly as
given and never substitutes a backend. Each entry reports whether the backend is
usable right now (installed, credentialed, permitted by policy, within limits),
the model type it would compile to, deterministic reason codes, blocking errors,
and the same warnings validate_optimization_problem would give for that backend.
Based only on the problem structure, backend capabilities and server policy —
no cost estimation, no network requests, no quota consumed.
A problem with integer variables adds reason codes: R_INTEGER_NATIVE when the
backend compiles to cqm and takes integers as they are, R_INTEGER_ENCODED when
it compiles to bqm and must binary-encode them, and R_INTEGER_BLOWUP when that
encoding also raises an INTEGER_QUADRATIC_BLOWUP warning — a backend carrying
that code is ranked after the ones without it.
The entries are ordered best first, so the first usable entry is the default
choice when the user named no backend. Among backends of the same kind, two
reason codes come from a backend's declared, measured structural preference:
R_DENSE_STRENGTH when it declares that it reaches the same energy as its
peers in a fraction of the time on large dense unconstrained models and this
problem is one, R_PENALTY_WEAKNESS when it declares a lower hit rate on
models whose hard constraints compile to penalties and this problem has an
effective hard constraint. Both shapes are bqm-path shapes, so a backend that
compiles to cqm is never matched. A backend carrying the first is ranked
ahead of its neighbours, one carrying the second behind them; neither changes
which backends are usable.
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| valid | Yes | The problem's own validity. When false, recommendations is empty. |
| errors | No | The problem's errors when it is invalid; empty otherwise. |
| advisory | No | A fixed sentence restating that a solve always uses problem.solver.backend as given and never substitutes one. |
| recommendations | No | Every registered backend, ranked best first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It states the tool makes no network requests, consumes no quota, and does no cost estimation. It explains the reason codes (R_INTEGER_NATIVE, R_INTEGER_ENCODED, R_INTEGER_BLOWUP) and the ranking logic (R_DENSE_STRENGTH, R_PENALTY_WEAKNESS) in detail, disclosing exactly what each entry reports and how ordering works. This is exemplary transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but each sentence adds substantive information: the advisory nature, the report contents, the integer reason codes, and the ranking rules. It is front-loaded with the core purpose and then systematically covers details. While not brief, it is well-structured and free of fluff, earning a score slightly above the median.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description is not required to explain return values, yet it still describes what each entry reports (usability, model type, reason codes, blocking errors, warnings) and explains the ordering rules. It covers all important behavioral aspects and edge cases (integer variables, dense unconstrained models, penalty weaknesses). The description is complete enough for an agent to understand when and how to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter, 'problem', which is a reference to a well-structured OptimizationProblem schema. The schema_description_coverage for the top-level parameter is 0%, so the description should compensate, but it adds no new information about the parameter beyond 'given a problem'. However, the schema itself provides exhaustive descriptions of all nested fields, so an agent can infer the parameter's meaning from the schema. The description does not add value here, but it does not mislead; a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, unambiguous statement: 'Rank the solver backends of this server for a given problem, without solving it.' It immediately distinguishes the tool from its siblings by explicitly noting that solve_optimization always uses the given backend and never substitutes, and by referencing validate_optimization_problem's warnings. This makes the tool's role and boundaries crystal clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the advisory nature and explicitly contrasts with solve_optimization, stating it never substitutes a backend. It also mentions it gives the same warnings as validate_optimization_problem, implying a validation-like use case. However, it does not explicitly state when NOT to use this tool or when to prefer one of the siblings beyond the advisory contrast, so it stops just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
solve_optimizationA
Solve a structured binary or bounded-integer combinatorial optimization problem.
Use this when the user asks which items to take within a budget, weight or
capacity, how to assign people or jobs to seats, shifts or machines, in
which order to visit a handful of places, or how to split things into
groups or pick a subset that meets several requirements at once — anything
expressible as yes/no or bounded-count decisions with a linear or quadratic
score and linear rules, even when the user never says "optimization".
Call this only after translating the user's request into explicit binary or
bounded-integer variables, an objective (linear/quadratic, minimize or
maximize), and hard or soft linear constraints. Do not pass natural-language
requirements.
An integer variable is declared with "type": "integer" plus integer
lower_bound and upper_bound (both required), and the problem must then carry
"version": "1.1" at its top level. A backend that compiles to bqm encodes
each integer in binary, so the compiled size grows with the range of the
bounds; a backend that compiles to cqm takes integers natively. Integer
values come back as ints inside their declared bounds.
Inequality constraints (<=, >=) require integer coefficients and right-hand
sides. Soft constraint weights are in objective units.
Leave solver.penalty_multiplier at its default unless a previous result was
infeasible on a remote backend; hard constraint penalties are managed by the
server. Call validate_optimization_problem first when planning to use a
remote backend or when the problem is large; on a local backend an invalid
document comes back as status invalid_problem with the same errors and
recommended_action, so solving directly spends nothing. Use
get_optimization_capabilities when you need the backend list or the full
schema. Use recommend_backend to compare backends; the choice remains yours.
Returns ranked feasible solutions with per-constraint evaluations, or a
structured error with a recommended_action. When the status is infeasible,
read infeasibility for the candidate that came closest to feasibility and
the share of candidates each hard constraint rejected, so the answer can
name the binding requirement instead of only reporting failure. Whatever
the status, the result's warnings are the same advisory warnings
validate_optimization_problem gives for this backend (an ignored seed or
parameter, a wide integer range, a negligible soft weight, ...) followed
by any raised during the run; read them before trusting a weaker answer
than expected. A field the schema does not declare is a tool error naming
its path, never ignored. If the user did not name a backend, say in the
answer which backend ran and why.
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| errors | No | Structured failures. Empty on success. |
| status | Yes | The single verdict for the request. success: at least one feasible solution was found and ranked. infeasible: the pipeline ran but no candidate satisfied every hard constraint under re-validation — check infeasibility_proven before concluding none exists. invalid_problem: the problem failed validation and no backend was invoked. solver_error: the backend failed while executing, or a remote vendor reported an error. backend_unavailable: the requested backend is not registered, not installed, disabled by policy or missing credentials — there is never a silent fallback. resource_limit_exceeded: a request parameter or the compiled size exceeded a server-side ceiling; values are refused, never clamped. configuration_error: the server's own configuration is wrong, not the problem. |
| backend | Yes | Registry name of the backend that ran, or null when the request failed before a backend was chosen. |
| message | No | Human-readable one-line summary of the result. On success: which backend produced it, whether optimality is proven, the rank-1 objective with its direction (and its soft violation when non-zero), how many distinct candidates the attempt saw, how many were feasible and how many are returned, and the attempt number when a retry produced it. On infeasible: why nothing feasible was found. On a failure: the first error's message. Deterministic — it never contains timings. Null when there is nothing to add. |
| attempts | Yes | One entry per compile/solve/validate attempt, in order. |
| metadata | No | Sanitized execution facts about the last completed attempt, local backends included. Null when the request failed before any solve finished; the service never invents it. |
| warnings | No | Non-blocking advice, same structure as an error: the warnings validate gives for this backend, then any raised during the run. Present whatever the status, except invalid_problem. |
| solutions | Yes | Ranked feasible solutions, best first, at most solver.top_k. Empty unless status is success. |
| elapsed_ms | No | Wall-clock milliseconds measured by the service from entering solve to returning, problem validation and any wait for a concurrency slot included. Unrelated to metadata.timing_us, which is what a vendor reports about its own side, and different on every run. |
| infeasibility | No | Why the last attempt found nothing feasible. Present only when status is infeasible and that attempt had candidates to diagnose; null on every other status, and on an attempt that received no samples at all. |
| optimality_proven | No | True only on success when an exhaustive backend enumerated every assignment: rank 1 is then the global optimum of ranking_score, not merely the best candidate seen. Always false on a heuristic or remote backend. |
| objective_direction | Yes | Echoed from the problem, so a consumer can interpret objective_value without re-reading the input. Null when the request failed before the problem was read. |
| annealbridge_version | No | The installed package version that produced this result; "unknown" outside an installed distribution. |
| infeasibility_proven | No | True only when an exhaustive backend actually enumerated every assignment and found none feasible. On a heuristic or remote backend an infeasible result only means 'not found under this configuration'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it discloses return behavior (ranked feasible solutions, structured errors, infeasibility details), warning semantics, backend-specific integer encoding differences, and how undeclared fields become tool errors. It also explains penalty_multiplier behavior and when to leave it at the default, which is beyond what the schema alone communicates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place—use cases, prerequisites, integer/version mechanics, constraint typing, validation guidance, sibling routing, and result interpretation are all packed in without fluff. It is front-loaded with purpose and examples before moving into technical and operational details, making it easy for an agent to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with a single large structured parameter, an output schema, and three sibling tools, the description is remarkably complete. It covers when to call, what to translate first, how to handle validation, what results mean, how to read infeasibility information, what warnings imply, and how to report backend choice. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite schema description coverage of 0%, the description compensates by explaining how to construct the problem parameter: explicit binary/bounded-integer variables, objective direction, linear/quadratic terms, and hard/soft constraints. It adds semantics not present in the schema, such as the version 1.1 requirement for integer variables, integrality rules for inequalities, soft-weight units, and integer encoding tradeoffs across backends. It does not enumerate every solver field, but it references get_optimization_capabilities for the full schema, so the gap is acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action—solving a structured binary or bounded-integer combinatorial optimization problem—and immediately grounds it in concrete user scenarios (budget, capacity, assignment, ordering, grouping). It also distinguishes itself from siblings by naming validate_optimization_problem, get_optimization_capabilities, and recommend_backend as separate tools, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool ('Use this when the user asks which items to take...'), when not to use it ('Do not pass natural-language requirements'), and what must be done first ('Call this only after translating the user's request into explicit... variables'). It also names sibling tools with clear conditions, such as validating first for remote backends, using get_optimization_capabilities for the backend list, and using recommend_backend to compare backends.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_optimization_problemA
Check a structured optimization problem without solving it.
Returns semantic errors (each with a recommended_action), advisory
warnings, and an estimate of the compiled variable count including slack
bits. Call this before solve_optimization when planning to use a remote
backend, so problems can be fixed before spending quota, or when the
problem is large (many variables, wide integer ranges). On a local backend
solve_optimization can be called directly: an invalid document returns
status invalid_problem with the same errors and recommended_action.
Nothing is compiled or solved and no network requests are made.
solve_optimization reports the same warnings for the same backend, so
skipping this call never hides them; calling it first only saves the
solve.
A field the schema does not declare is a tool error naming its path,
never ignored: check the schema (get_optimization_capabilities with
include_schema: true returns it as problem_json_schema) before inventing
one.
The estimate follows the model type the chosen backend compiles to. On a
bqm backend it counts the slack bits of every inequality constraint plus
the binary-encoding bits of every integer variable, so a wider
lower_bound..upper_bound range costs more compiled variables. On a cqm
backend integer variables are native and no encoding bits are counted.
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| valid | Yes | Decided by errors alone. Warnings are advisory and never make a problem invalid. |
| errors | No | Every error found in one pass — validation does not stop at the first. Empty means the problem is safe to compile. |
| warnings | No | Advisory findings that do not block a solve. Only produced when there are no errors. |
| model_type | No | Which compiler path the estimate assumed. Null when the problem is invalid. |
| objective_scale | No | The upper bound on the objective's range, used to size hard penalties and to judge whether a soft weight is meaningful. Null when the problem is invalid. |
| estimated_compiled_variables | No | Compiled size without building a model: on the BQM path, binary variables + integer-encoding bits + slack bits; on the CQM path, variables + integer slacks. Null when the problem is invalid. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and meets it: it states that nothing is compiled or solved, no network requests are made, and that solve_optimization reports the same warnings. It also discloses the strict unknown-field behavior, which is exactly the kind of non-obvious trait an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: return contents, usage conditions, side-effect guarantees, schema-checking advice, and estimate semantics are all packed in without repetition. The most important scoping information, 'without solving it,' is front-loaded in the first sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers return contents, when to call versus alternatives, side effects, unknown-field policy, and backend-dependent estimation semantics. Since an output schema exists, the shape of return values need not be restated. Nothing an agent needs to decide whether and how to invoke this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'problem' parameter has no top-level schema description (0% coverage), so the description must compensate. It does add meaningful behavior: undeclared fields become tool errors, integer range width affects the compiled-variable estimate, and bqm/cqm backends count variables differently. However, it does not summarize the required structure of the problem object itself, leaving that to the embedded $ref definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check a structured optimization problem without solving it,' and then names what the tool returns (semantic errors, advisory warnings, compiled-variable estimate). It clearly differentiates itself from the sibling solve_optimization by framing this as a pre-solve validation step for remote backends.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: call before solve_optimization for remote backends or large problems, and skip it on local backends where solve_optimization returns the same errors directly. It even states that skipping this call never hides warnings, so the agent knows exactly what is and is not gained by calling it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
get_optimization_capabilities - First observed
recommend_backend - First observed
solve_optimization - First observed
validate_optimization_problem
TDQS
Scored across 4 tools
Each tool has a clearly distinct role: validating, solving, querying capabilities, and ranking backends. Even though validate_optimization_problem and solve_optimization both report warnings, their purposes are explicitly separated and cross-referenced, so an agent should not confuse them.
All tool names follow a consistent snake_case verb_noun pattern: validate_optimization_problem, recommend_backend, get_optimization_capabilities, solve_optimization. The verbs are descriptive and uniform in style.
Four tools is a well-scoped set for this server's purpose: validate before solving, solve the main task, inspect capabilities, and get backend recommendations. Each tool earns its place without redundancy or bloat.
The tool surface covers the full workflow of an optimization server: understanding accepted problems, validating them, choosing a backend, and solving. No obvious dead ends or missing core operations are apparent for the stated domain.
Maintenance
Related MCP Connectors
Pay-per-use tool marketplace for AI agents. Search, price-check, and call APIs via MCP.
Give AI agents identity, scoped access, trusted context, and verifiable actions through MCP.
One MCP tool for verified AI-agent outcomes with success-only charging.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceMCP-ORTools integrates Google's OR-Tools constraint programming solver with Large Language Models through the MCP, enabling AI models to: Submit and validate constraint models Set model parameters Solve constraint satisfaction and optimization problems Retrieve and analyze solution21MIT
- AlicenseNot gradedqualityDmaintenanceEnables solving Constraint Satisfaction Problems (CSP) like N-Queens, graph coloring, and Sudoku, as well as Linear Programming optimization problems through both MCP tools and HTTP API endpoints.2MIT
- AlicenseCqualityCmaintenanceEnables Large Language Models to submit and solve constraint satisfaction and optimization problems using Google OR-Tools through JSON model specification.11MIT
- FlicenseBqualityDmaintenanceEnables AI assistants to solve linear, integer, mixed-integer, and knapsack optimization problems using Google OR-Tools via a simple JSON interface.3-