MaterialsCodeGraph
Server Details
Materials simulations for AI agents over MCP: binding token quote before every run, record after it.
- Status
- Healthy
- Uptime
- 13.6% over 23 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 12 tools
Most tools target distinct lifecycle stages (discovery via mcg_catalog, pricing via mcg_quote/mcg_rates, launching via mcg_run, drafting via mcg_stage, cancellation, run/study inspection, artifact listing vs. content fetch). The main overlaps are mcg_quote vs. mcg_rates and mcg_fetch_artifacts vs. mcg_get_artifact, but the descriptions clearly delineate manifest vs. single-file and price-binding vs. public rate card.
All tools share a consistent mcg_ prefix and snake_case naming, mostly with verb_noun forms (cancel_run, get_run, list_runs, fetch_artifacts, get_artifact). A few are noun-only resource identifiers (catalog, clusters, rates) which is a minor deviation but still readable and predictable.
Twelve tools is well-scoped for a simulation-quoting and running platform with a full lifecycle (discovery, quoting, staging, launching, cancelling, polling, artifact access, rate card). Each tool earns its place without obvious redundancy.
Core CRUD/lifecycle coverage is strong: discovery, cluster targeting, quoting, draft staging with uploads, launch, cancel, run/study inspection, artifact manifest and content fetch, and rate card. Minor gaps include no tool for cluster creation (only via POST /api/clusters) and no in-agent binary artifact retrieval, but these are noted and workable.
Available Tools
12 toolsmcg_cancel_runADestructiveInspect
Cancel a run from any non-terminal state before collecting (a collecting run's outputs are delivered and charged: 409). The cancel request is recorded immediately and returns run_id and cancel_requested true; the platform kills the machine and terminalizes the run as state killed on its next orchestration tick (about 60 s). Poll mcg_get_run to observe the terminal state.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | The run id to cancel. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=true, but the description goes far beyond: it discloses that the request is recorded immediately, returns run_id and cancel_requested=true, that the machine is killed and the run terminalizes as 'killed' on the next orchestration tick (~60 s), and the exact 409 failure condition for collecting runs. That is unusually rich behavioral context for an irreversible operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler, front-loaded with the cancellation scope before the error case and the async lifecycle. Every clause carries operational meaning (states, 409, timing, polling).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema, the description supplies everything an agent needs: eligible states, the blocking error, the asynchronous completion model with timing, the returned fields, and the polling tool. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single run_id parameter, and the schema already describes it fully, so the baseline 3 applies. The description references run_id only in the return-value context, adding no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Cancel a run') and immediately scopes it to 'any non-terminal state before collecting', which distinguishes it from read/collection siblings. An agent can tell exactly what state transition this performs without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (non-terminal states), a when-not with the consequence spelled out ('a collecting run's outputs are delivered and charged: 409'), and names the follow-up tool (mcg_get_run) for observing the result. Routes the agent correctly against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_catalogARead-onlyInspect
Enabled simulation kinds (each with example mcg_quote arguments that price, or a note naming what must exist first; traced says whether the kind's default lineage returns an OpenMaterials record, traced_note why or why not), the costed lineage menu (runnability, physical conditions, indicative token costs), the external-frontier graphs mcg_quote accepts, recipes (pass a recipe's quote object to mcg_quote unchanged), and codes (versions, licenses, pins; the only software a run may execute). Call before mcg_quote. A kind or code absent here cannot run.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already declaring safety, the description adds substantial behavioral context: it is a prerequisite catalog, it lists what data is available, and it constrains runs to kinds and codes present here. It does not discuss freshness, caching, or rate limits, but those are minor for this read-only catalog.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the catalog's contents and then closes with two short usage sentences. The first sentence is long and dense with parentheticals, but every clause contributes distinct information about what the catalog contains and how it relates to mcg_quote.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters, no output schema, and a readOnly annotation, the description gives a complete enough picture of what the tool returns and why it matters. It covers the catalog's sections, the prerequisite relationship to mcg_quote, and the consequence of a kind or code being absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there is no parameter semantics to explain. The rubric's baseline for zero parameters is 4, and the description appropriately does not add irrelevant parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates the catalog's contents (enabled simulation kinds, costed lineage menu, external-frontier graphs, recipes, codes) and distinguishes its role from mcg_quote by stating it must be called first. However, it does not open with a specific action verb such as 'Returns' or 'Lists', so it falls just short of the highest clarity tier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call before mcg_quote' and adds 'A kind or code absent here cannot run,' making both when-to-use and when-not conditions clear. It also distinguishes the recipe path by telling the user to pass a recipe's quote object to mcg_quote unchanged.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_clustersARead-onlyInspect
Connected clusters: id, name, preset (hive or auto), status (pending until the agent's first sync; connected; stale: the agent stopped syncing; revoked; scheduler_down: the agent syncs but cannot reach the scheduler, nothing is offered; agent_outdated: an auto cluster's agent sends no inventory; snapshot_rejected: its last snapshot was over the size cap; snapshot_stale: no fresh capacity snapshot, nothing is offered; nodes_unavailable: none of an auto cluster's nodes accept jobs (drained, down, in maintenance or not responding), nothing is offered), status_detail (why, when not connected), classes (resource classes offered per partition, with per-job GPU, CPU and memory maxima and the GRES type), unclassified_gpus (GPU types present but not offered: type, count, reason), untyped_gpus, last sync time. Pass target {cluster_id} to mcg_quote to run there at zero hardware cost; only a connected cluster quotes. Clusters are created with POST /api/clusters (on the machine host with a spend-scoped key).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behavior: what 'pending', 'stale', 'scheduler_down', 'agent_outdated', 'snapshot_rejected', 'snapshot_stale' and 'nodes_unavailable' each mean, and specifically which states result in 'nothing is offered'. That is substantial operational context an agent cannot get from the annotation alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the returned fields and the status semantics are genuinely useful, but the content is delivered as a single dense run-on sentence with a very long nested parenthetical, which hurts scanability. It earns most of its length but could be broken into structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by enumerating return fields and status meanings, and it notes cluster creation via POST /api/clusters. For a zero-parameter read tool this is quite complete; only a plain statement that this is a listing call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The only parameter-like reference (target {cluster_id}) belongs to mcg_quote, not this tool, so there is no parameter semantics to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly establishes the resource (connected clusters) and enumerates the fields returned (id, name, preset, status, classes, etc.), making the tool's purpose evident as a cluster listing/inspection call. It does not state the verb explicitly and differentiates only loosely from siblings like mcg_catalog, but the resource scope is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides one concrete downstream linkage – 'Pass target {cluster_id} to mcg_quote to run there at zero hardware cost; only a connected cluster quotes' – which implies when this tool matters. However, it never states when to call this versus other list/inspection siblings, and offers no explicit exclusions, so guidance is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_fetch_artifactsARead-onlyInspect
List a run's collected artifacts as a typed manifest: artifacts, a list of path, bytes, stage and sha256 per file, plus removed on a file the retention sweep deleted (its bytes are gone; do not fetch it); truncated; source. Never returns file contents; fetch one file with mcg_get_artifact.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | The run id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint, so the description does the heavy lifting: it discloses the manifest field set, that retention-swept files appear as 'removed' with bytes gone, and that output may be truncated. Auth requirements and rate limits are not covered, but for a read-only listing this is strong disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the return shape, with no filler sentences. The semicolon-chained field list is dense but each clause carries real information, so it reads as compact rather than bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the manifest fields (path, bytes, stage, sha256, removed, truncated, source) and the caveat on deleted files. Nothing an agent needs to interpret or correctly call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is a single required run_id, so the schema already carries the semantics. The description only references 'a run's' artifacts and adds no format or id-syntax detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List a run's collected artifacts as a typed manifest') and immediately distinguishes itself from the sibling that returns file contents. An agent can separate this from mcg_get_artifact without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when not to use it ('Never returns file contents') and names the alternative ('fetch one file with mcg_get_artifact'), plus a second exclusion ('removed ... do not fetch it'). Routing conditions are fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_get_artifactARead-onlyInspect
Fetch one artifact file's TEXT content, up to max_bytes (default 65536), with an explicit truncation note. Binary content REJECTS with unsupported_content; fetch binaries from the run's public page.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | The artifact sub-path, e.g. results/report.md. | |
| run_id | Yes | The run id. | |
| max_bytes | No | Max bytes to inline (default 65536). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations carry the read-only safety profile, and the description adds genuinely beyond it: truncation at max_bytes with an explicit truncation note, and a named failure mode (unsupported_content) for binary input. It stops short of auth requirements or how truncation interacts with large files, but the disclosed behavior is well above the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences with no filler; the primary action and its constraint come first, with the failure mode and escape hatch immediately after. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still tells the agent what it gets back (text content, possibly truncated with a note) and what errors to expect. Missing only edge details such as behavior on nonexistent paths or very large max_bytes, which are minor for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so path, run_id, and max_bytes are already documented with types and the 65536 default. The description merely restates the max_bytes default and the text-only constraint rather than adding format or syntax guidance, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ("Fetch one artifact file's TEXT content") with a clearly bounded scope: a single file, text only. The words "one" and "TEXT" implicitly separate it from the sibling mcg_fetch_artifacts (plural) and from binary retrieval, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-not rule (binary content REJECTS) and points to an alternative destination ("the run's public page") for binaries. It does not explicitly contrast with the sibling mcg_fetch_artifacts for multi-file retrieval, so usage is clear but not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_get_runRead-onlyInspect
Fetch a run's state, event timeline and LLM spend. The payload leads with a derived summary (state, current_stage, billing {quoted_tokens, charged_tokens, hold, policy}, fault, refusal, next, last event) for cheap polling, then app_url (the run page, for the human), study and headline. events holds the newest events that fit about 12 KB plus the last complete or fail event; events_omitted counts the others. In a payload over 2 KB, each value over 256 chars becomes {omitted_bytes, see}, where see names the route that returns the full event. A report body_md over 6,000 chars becomes the same stub; a brief body_md is cut at 6,000 chars and ends in an ellipsis; llm_spend.calls holds the newest 10 calls, and calls_omitted counts the rest. A done, node-addressed run also carries record (its OpenMaterials SimulationRecord), openmaterials_url (the datasheet link) and openmaterials_label: null, or the commons' note when the record cites a model or configuration that is not in the public registry (record.unregistered lists each), so its values cannot be checked against the commons.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | The run id from mcg_run. |
mcg_get_studyARead-onlyInspect
Fetch a study (the project a quote's study named) with its member simulations and their latest runs. Pass study_id or the exact name. Returns study, members (each with app_url) and app_url.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | The exact study name given to mcg_quote or mcg_stage. | |
| study_id | No | The study id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes the safe-read profile, and the description adds useful detail about the payload shape (members, each with app_url, plus app_url and latest runs). However it discloses nothing about missing studies, name-vs-id mismatch behavior, or result size limits, so it adds only moderate value beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the return payload front-loaded after the verb, no filler. The phrasing 'the project a quote's `study` named' is slightly contorted but still compact and information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately summarizes the return shape (study, members, app_url), and both optional params are covered by the schema. Missing only edge-case behavior (unknown name/id, mutual-exclusion precedence), which is a minor gap for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both params are optional, so the schema already documents them. The description adds that `name` must be the exact name given to mcg_quote/mcg_stage and frames the two as alternatives, which is mild extra meaning but not a full explanation of precedence when both are supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (study), and enumerates what it returns: member simulations and their latest runs. It also ties the resource back to mcg_quote/mcg_stage ('the project a quote's `study` named'), which helps an agent recognize the domain object, though it does not explicitly distinguish itself from siblings like mcg_get_run or mcg_list_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives parameter-level advice ('Pass study_id or the exact name') but no when-to-use/when-not guidance relative to alternatives such as mcg_get_run or mcg_list_runs. Usage is implied by the link to the study named in mcg_quote/mcg_stage rather than stated as a selection rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_list_runsARead-onlyInspect
List runs, newest first: every run of this key's owner, or one quote's runs when quote_id is given. Each row carries state, spend_tokens, created_by_key_name (null when launched in the browser), study, headline (the AI brief's first line) and app_url. Returns runs.
| Name | Required | Description | Default |
|---|---|---|---|
| quote_id | No | Optional: the quote_id (simulation) to list runs for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers safety. The description adds useful behavioral context: results are ordered newest first, each row contains specific fields, and created_by_key_name is null for browser-launched runs. It does not disclose pagination or result limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary behavior and efficiently lists return fields. The trailing sentence 'Returns runs.' is redundant with the opening verb and slightly weakens conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one optional parameter and no output schema, the description adequately explains scope, ordering, and returned fields. Missing pagination/limit details and sibling-tool routing keep it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents quote_id, so baseline is 3. The description adds meaning by clarifying the default scope when quote_id is omitted ('every run of this key's owner') and the filtered scope when it is given ('one quote's runs').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('List runs') and states the scope: all runs for the key's owner, or one quote's runs when quote_id is given. It does not explicitly differentiate from siblings like mcg_get_run, but the list-vs-get distinction is clear from the tool name and description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly describes the two modes: default all runs of the owner, or filtered to one quote's runs when quote_id is provided. There is no explicit guidance on when to choose this over alternatives, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_quoteAInspect
Price a simulation. FREE: no hardware starts. Returns quote_id, quote_hash, tokens, rate_card_version, wall_s, expires_at, the plan, and record: {traced: true, node, node_uid, map_version (the openmaterials graph version the node is pinned at; the commons map version is published at openmaterials.ai/data/version.json), material (the material the pinned conditions name: what the run computes)} when the completed run's OpenMaterials record is keyed on that public map node (measured instances cite the method that ran), else {traced: false, reason}; read it before you spend. The price is a binding ceiling: never more than this. mcg_run accepts only this quote_id with this quote_hash before expires_at and REJECTS a stale hash (quote_mismatch) or an expired quote (quote_expired). Arms, exactly one: kind from mcg_catalog (spec passed through verbatim, spec.params set the price, engine overrides the default only with a known engine id); external_frontier, an OpenMaterials external-solve envelope from mcg_catalog's graphs (also returns request_id, plan_digest and the compiled plan; an envelope not byte-equal to a supported graph REJECTS with unsupported_graph); draft_id (see the parameter). Also returns app_url, the owner's page in the app (for the human, never fetched by the agent) and, on the kind arm, warnings: the inputs the app cannot render without (material or structure_xyz, study); pass them. accuracy_study needs no hardware_contract up front: the reply's contract_source says whether the run binds this simulation's published contract (published) or the labeled default synthetic one (default_synthetic).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Simulation kind from mcg_catalog, e.g. bulk_kappa. | |
| name | Yes | A human label for the simulation. | |
| spec | No | The spec object, e.g. {params: {...}, spec_yaml: '...'}. params drive the price. interface_tbc: params.production_steps and params.equilibration_steps set the NEMD budget (the run total across all interfaces) within the bounds the reply's certification.settable states (a refused run's summary.next names the key to raise); the price follows the budget. | |
| study | No | The study (project) this simulation belongs to, by exact name; created if absent. Every quote for one investigation should name the same study so the app groups them together. | |
| engine | No | Optional engine id override (rejected if unknown). | |
| target | No | Run on one of your connected clusters (see mcg_clusters) instead of the cloud: hardware cost 0, gate fee per the rate card and plan, wall time includes the scheduler's queue-wait estimate. The cluster's status must be connected. | |
| draft_id | No | Third arm: price a draft that mcg_stage created and whose inputs have been uploaded (refuses with inputs_missing while any declared input is absent). Also re-prices an already-quoted simulation, after its quote expired or to reproduce a run (pass the run's quote_id, its simulation_id): it clones the spec, the map pin and the staged inputs (not contract files, nor structure files the engine does not read; more than 25 inputs refuse) and returns the clone's fresh quote_id (reply carries requoted_from, and rerun_notes naming skipped inputs and structure overrides). Mutually exclusive with every other argument but name and target. | |
| material | No | Kind arm only: the material or formula being simulated (e.g. Si, MoS2). Also selects the reference structure. On a traced run whose engine fixes its own conditions it must match the spec's material. | |
| conditions | No | Kind arm only: physical conditions as key/value pairs, e.g. {T: 300}. Values are numbers, strings or lists of up to 16 of them, except on an adapter engine, which runs them: there any value its conditions schema allows (lists, booleans, null) over the keys it declares, judged at quote. On a traced run whose engine fixes its own conditions, a key or value its record does not hold is refused; the error names what to send or omit. | |
| max_tokens | No | Graph arm only: quote-time hard ceiling; a quote above it REJECTS (never clamps). | |
| structure_xyz | No | Kind arm only: the simulated structure as extended-XYZ text; staged as the input structure.xyz so the app renders the exact geometry (takes precedence over the material reference). | |
| external_frontier | No | The OpenMaterials #110 external-solve envelope, verbatim. Mutually exclusive with kind/spec/engine. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description carries the full burden. It discloses that the call is FREE, lists returned fields (quote_id, quote_hash, tokens, plan, record tracing, app_url, warnings), states binding ceiling semantics, and names rejection reasons (quote_mismatch, quote_expired, unsupported_graph). Exceptionally thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense single paragraph with many clauses and no bullets or clear sections. Every sentence is informative, but the lack of structure makes it hard to parse quickly. Length is justified by complexity but could be better organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 12 parameters, nested objects, no output schema, and a complex multi-arm workflow, the description covers purpose, usage, return values, error conditions, and workflow constraints (e.g., mcg_run acceptance). It is complete enough for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds key semantics beyond the schema: mutual exclusivity of the three arms, that spec.params drive the price, and draft_id re-pricing and clone behavior. These clarify parameter interactions meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Price a simulation.' It clearly distinguishes from siblings like mcg_run (which runs) and mcg_catalog (which supplies kinds). The purpose is unambiguous and front-loaded.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'read it before you spend' and details three mutually exclusive arms (kind, external_frontier, draft_id) with their conditions, including when to use each and what happens on rejection. Guides the agent on selection and workflow integration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_ratesARead-onlyInspect
The public rate card: token_usd, overheads, rate_card_version. FREE, no arguments; also a connectivity check.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds genuinely new behavioral context: that the call is FREE (no billing/quota impact) and requires no arguments, plus that it doubles as a connectivity probe.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse fragments with zero filler, and the core purpose is front-loaded. The telegraphic style is efficient, though the fragmentary phrasing is slightly abrupt even for a trivial tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-arg, read-only lookup with no output schema, naming the returned fields and the free/connectivity nature covers what an agent needs. Only auth requirements and the distinction from mcg_quote are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description usefully notes 'no arguments' and lists the fields the caller can expect back, which partially compensates for the absence of an output schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('the public rate card') and enumerates the values it exposes (token_usd, overheads, rate_card_version). It is distinguishable from most siblings, though it does not explicitly contrast with mcg_quote, the other pricing-related tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The dual-use hint ('also a connectivity check') and 'FREE, no arguments' imply when to reach for it, but there is no explicit when-to-use vs mcg_quote or statement of prerequisites/preconditions. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_runAInspect
Launch the quoted simulation. SPENDS TOKENS: provisions hardware and bills the account. Requires the quote_id and quote_hash from mcg_quote (a stale hash or an expired quote REJECTS: re-quote and retry) and a client-generated idempotency_key (replaying it returns the existing run instead of launching twice). Optional max_tokens is a hard ceiling: a quote above it REJECTS the launch, never clamps. Returns run_id.
| Name | Required | Description | Default |
|---|---|---|---|
| quote_id | Yes | The quote_id returned by mcg_quote. | |
| max_tokens | No | Hard token ceiling: the launch is REJECTED if the quote exceeds it. | |
| quote_hash | Yes | The quote_hash mcg_quote returned: proves the launch prices the exact quote the client saw. | |
| idempotency_key | Yes | Client-generated key; retrying a launch with the same key must not double-charge. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare idempotentHint=false, but the description asserts the opposite behavior: 'replaying it returns the existing run instead of launching twice' — i.e., repeated calls with the same arguments have no additional effect, which is exactly what idempotentHint=true means. This is a direct mismatch between the stated behavior and the structured hint. The description's other disclosures (token spend, billing, hard-ceiling rejection, staleness rejection) are valuable, but the idempotency conflict triggers the contradiction rule.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with the action and the cost warning ('SPENDS TOKENS') before prerequisites and edge cases. Every clause carries operational information; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter launch tool with no output schema, the description covers cost, required inputs, rejection conditions, retry path, idempotency behavior, and the return value ('Returns run_id'). An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents quote_id, quote_hash, idempotency_key, and the max_tokens hard ceiling; the description largely restates those (stale hash, no clamping, replay protection). The only marginal addition is the 're-quote and retry' workflow for rejected quotes. Baseline 3 is appropriate when the schema carries parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Launch the quoted simulation') and clearly positions the tool at the end of the mcg_quote -> mcg_run workflow. An agent can distinguish it from mcg_quote, mcg_get_run, and mcg_list_runs without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit prerequisites (quote_id and quote_hash from mcg_quote) and the recovery path for a stale/expired quote ('re-quote and retry'), plus the replay semantics for idempotency_key. It stops short of naming alternative siblings or stating when NOT to launch (e.g., cancel/observe via mcg_get_run), so it is clear context rather than full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcg_stageAInspect
Create a DRAFT simulation whose input files are uploaded before pricing. FREE. Takes the same kind, spec, study, material and conditions as mcg_quote plus inputs: one entry per file the run needs, each with a name and, for small text files, the text inline (staged immediately). Entries without text come back in uploads as presigned PUT URLs (valid 15 minutes, 4 GiB per file): upload each with a plain HTTP PUT of the file bytes, no key or header. Then call mcg_quote with draft_id to price; it REJECTS with inputs_missing until every declared input has arrived.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Simulation kind from mcg_catalog. | |
| name | Yes | A human label for the simulation. | |
| spec | Yes | The spec object; params drive the eventual price. | |
| study | No | The study (project) this simulation joins, by exact name; created if absent. | |
| engine | No | Optional engine id override. | |
| inputs | Yes | The input files the run needs. | |
| material | No | The material or formula simulated. | |
| conditions | No | Physical conditions: numbers, strings or lists of up to 16 of them (on an adapter engine, any value its conditions schema allows). On a traced run whose engine fixes its own conditions, a key or value its record does not hold is refused; the error names what to send or omit. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the whole burden and does so richly: cost ('FREE'), draft (non-final) state, upload mechanics (presigned PUT, 15-minute validity, 4 GiB per file, no key or header), and downstream rejection behavior. It discloses auth requirements, size/time limits, and failure mode beyond anything in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the FREE/draft scoping are front-loaded, and the equivalence to mcg_quote's parameters is stated compactly. It is dense and somewhat run-on in the later sentences, but nearly every clause (validity window, size cap, no-auth note) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description explains the return shape (uploads with presigned PUT URLs) and the end-to-end flow to pricing, including the failure mode. For an 8-parameter tool with nested objects, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds genuine meaning: it explains that an entry's text is staged immediately while an entry without text yields a presigned URL, and it links draft_id to the follow-up mcg_quote call. That goes beyond the schema's field-level wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Create a DRAFT simulation') and immediately scopes it against the sibling pricing tool by saying it takes the same kind/spec/study/material/conditions as mcg_quote plus inputs. An agent can distinguish mcg_stage (staging) from mcg_quote (pricing) without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the full workflow explicitly: stage to get upload URLs, HTTP PUT each file, then call mcg_quote with draft_id to price. It names the alternative tool and the exact condition that transitions to it, plus the precondition ('REJECTS with inputs_missing until every declared input has arrived').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
- First observed
mcg_cancel_run - First observed
mcg_catalog - First observed
mcg_clusters - First observed
mcg_fetch_artifacts - First observed
mcg_get_artifact - First observed
mcg_get_run - First observed
mcg_get_study - First observed
mcg_list_runs - First observed
mcg_quote - First observed
mcg_rates - First observed
mcg_run - First observed
mcg_stage
Related MCP Connectors
Agent Token Budget MCP — hard per-session token + spend cap with signed budget-exhausted
Paid token risk and security intelligence for AI agents over MCP with x402 payments.
- llm-busOAuthcom.llm-bus
Coordinate multiple AI agents over MCP: atomic claims, leases, shared ledger, handoffs, tasks.
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Related MCP Servers
- AlicenseAqualityDmaintenanceAn interactive MCP server for FAIRChem and ASE simulations that allows LLM agents to load a model once and steer relaxations, MD, NEB, phonons, and minima searches mid-flight, with live monitoring and code introspection.171MIT
- AlicenseAqualityCmaintenanceAn MCP server for computational materials science that enables AI assistants to generate LAMMPS simulations, parse outputs, analyze nematic order, detect plastic rearrangements, and estimate viscosity through 12 specialized tools.121MIT
- AlicenseNot gradedqualityCmaintenanceMCP-native scientific skills for reproducible computational biology and AI-driven drug-discovery workflows. It combines deterministic scientific tools with an MCP server to give AI agents real computational capabilities.Apache 2.0
- AlicenseAqualityAmaintenanceThe official MCP interface for AgentFEM, giving AI agents seven typed tools to create, validate, run, inspect, and verify finite-element simulations while preserving scientific evidence.720 PyPIApache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.