RQM Jobs MCP
Server Details
Delegate 60 specialist work products: 20 WaveEngine, 20 RQM Studio, and 20 RQM Robotics jobs.
- Status
- Healthy
- Uptime
- 54.0% over 43 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 38 tools
Most tools have distinct targets, and the robotics_* functions are well separated by name and description. However, boundaries are fuzzy among run_account_job / run_buyer_job / submit_*_job / settle_buyer_job, and the several get_job_* / get_account_job read tools could cause misselection without careful reading.
The set is almost entirely lowercase snake_case with a clear verb_noun structure, and the robotics_*_v1 prefix pattern is consistent. Minor deviations include run_ vs submit_ for similar submission actions and robotics_frame_convention_validation_v1 lacking the verb-initial form used by its siblings.
At 38 tools, the server is well past the practical 3–15 tool range and even exceeds the 25-tool threshold for 'too many.' The large cluster of narrow robotics_* analysis functions makes the surface feel bloated for an agent to navigate efficiently.
The job lifecycle is broadly covered: quote, fund, submit/run, cancel, monitor, retrieve results/artifacts, list, and recover by idempotency key. Minor gaps exist around explicit buyer-job status monitoring and post-funding confirmation, but these do not create dead ends for core workflows.
Available Tools
38 toolscancel_jobCancel an RQM jobAIdempotentInspect
Request cancellation of one authorized product job.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the key traits: readOnlyHint=false, idempotentHint=true, destructiveHint=false. The description adds the important nuance that this is a request, implying cancellation may not be immediate or guaranteed, and that only authorized jobs can be cancelled. It stops short of stating what happens to in-flight jobs or partial results, so it adds value but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence stating the action and scope. Every word earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. Annotations and the full schema cover the structured details. The only remaining gap is the absence of prerequisite or post-call guidance, but the definition is largely complete for a single-purpose cancellation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the anyOf variants fully document job_id, product, schema_version, and idempotency_key constraints. The description adds no further parameter meaning beyond what the schema already provides, which fits the baseline 3 when structured data does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Request cancellation) and resource (product job). It uses 'one authorized product job' to scope it, which differentiates from submission siblings like submit_quantum_job and run_buyer_job, though it doesn't name the primary alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: request cancellation for one authorized product job. There is no explicit when-to-use guidance or mention of alternatives such as checking status first with get_job_status, nor any warning about jobs already completed. Adequate but with clear gaps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_funding_sessionCreate an RQM MCP funding sessionCIdempotentInspect
Create an idempotent Stripe Checkout URL for human prepaid funding.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotent, non-destructive, open-world, non-read-only, yet the description repeats 'idempotent' rather than adding new behavioural detail. It omits what the returned URL does, expiry, and that funding is an external side effect on Stripe – useful context for an open-world mutation that the description fails to supply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with no filler and the key verb('create') front-loaded. Slight inefficiency because 'idempotent' duplicates the annotation rather than adding information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return structure need not be explained, and annotations cover safety flags. However, for a three-param, open-world mutation with zero schema descriptions, an agent needs to know what happens to a reused idempotency_key and what amount units mean – neither is addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description names zero of the three required parameters. The 'idempotent' claim loosely maps to idempotency_key but gives no syntax or collision semantics; amount and schema_version are wholly unmentioned in prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a ... Stripe Checkout URL') and adds the qualifying scope 'for human prepaid funding', which no sibling covers. It stops short of 5 only because it doesn't name or contrast with any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or prerequisite guidance. An agent cannot tell from the description whether this precedes run_buyer_job or is the sole funding path. 'Human prepaid funding' hints at a context but is not actionable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_account_balanceGet RQM MCP account balanceARead-onlyIdempotentInspect
Read the authenticated canonical Account Core prepaid USD balance.
| Name | Required | Description | Default |
|---|---|---|---|
| schema_version | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds two useful facts beyond them: the call is authenticated and the balance is prepaid and denominated in USD.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every word (authenticated, canonical, prepaid, USD) narrows the scope of the resource, so nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value structure need not be described, and the annotations fully cover safety and idempotency. The description is adequate for a no-argument read tool; only a hint about the account scope (e.g. per-workspace vs global) is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter (schema_version) is an opaque version const that the description never mentions. However, a fixed const discriminator carries no real semantic decisions for the agent, so the omission is minor; baseline 3 holds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ("Read") with a specific resource ("authenticated canonical Account Core prepaid USD balance"), and no sibling tool competes for that purpose. It is clear but relies on internal jargon ("canonical Account Core") that an agent has no external referent for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when/when-not guidance or named alternatives, but usage is strongly implied by the name and the read-only nature of a balance check. An agent can infer this is the tool for checking prepaid funds, though the description never says so.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_account_jobRecover an account-funded jobARead-onlyInspect
Read this agent's saved job, result and receipt. Never authorizes another purchase.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it readOnly and closed-world; the description reinforces this by saying 'Read' and adds a specific safety guarantee: it never authorizes another purchase. This is valuable context because the title 'Recover' could otherwise hint at re-running a paid job. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the main action front-loaded and the safety caveat after. Every word adds information; no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only getter, the description covers what is returned (job, result, receipt), whose jobs are in scope (this agent's), and the side-effect profile (no purchase). It does not describe response shape or error cases, but with readOnly annotations and a UUID parameter, the remaining gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, and it does supply scope: job_id refers to a saved account-funded job belonging to this agent. It does not explicitly explain how to obtain job_id or what values are valid beyond the UUID format. For a single self-descriptive parameter this is adequate but minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb ('Read') and resource ('this agent's saved job, result and receipt'), clearly scoping the operation. The title and 'account-funded' language distinguish it from buyer-job and submission tools. This is unambiguous relative to siblings like get_job_status and get_job_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear retrieval context: this tool is for reading a saved account-funded job's job, result, and receipt for the current agent. 'Never authorizes another purchase' is an explicit when-not that warns against using it for purchase or re-run workflows. It does not name alternative sibling tools, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_capabilityGet one RQM job capabilityBRead-onlyIdempotentInspect
Return one full approved public Wave capability descriptor.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds real context with 'approved public' (only approved, publicly visible capabilities are returned) and 'full' (complete descriptor rather than a summary), but says nothing about auth, error behavior, or not-found handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the verb and resource front-loaded and no filler. It is efficient, though arguably too terse for a tool whose schema exposes four request variants.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value description is not required. Against a four-variant anyOf request schema and a large sibling set, however, the description omits variant selection guidance (which schema_version/product to send) and any routing to list_capabilities, leaving real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Per the context there are 0 top-level parameters and schema description coverage is 100%, which normally sets a baseline of 4. The description adds no meaning of its own and does not explain the product/capability_id/schema_version request variants (v1-v4) expressed in the anyOf; its 'Wave' wording could mislead about the product field, so it sits below the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and resource ('one full approved public Wave capability descriptor'), and the singular 'one' implicitly contrasts with the sibling list_capabilities. However it never names that sibling explicitly, and narrowing to 'Wave' is slightly at odds with the schema's multi-product (wave/quantum/robotics) request variants.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use statement, no mention of list_capabilities as the alternative for enumerating capabilities, and no prerequisites or conditions. The agent must infer that this is the single-item lookup variant purely from the word 'one'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_artifactGet one Wave job artifactCRead-onlyIdempotentInspect
Retrieve one principal-scoped bounded Wave result artifact.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| product | Yes | ||
| artifact_id | Yes | ||
| schema_version | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds 'principal-scoped' (an authorization scoping constraint) and 'bounded' (result-size constraint), which is genuine extra context, but it says nothing about error behavior when the artifact is missing or still pending.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with no filler, which is good front-loading, but it is terse to the point of under-specification rather than being an efficient complete statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but with four required, wholly undocumented parameters and no usage context, an agent cannot tell how to obtain a valid artifact_id or how this reads contextually against get_job_result and get_job_status.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description supplies no parameter meaning whatsoever. The artifact_id's required 32-character hex pattern and the fixed schema_version/product constants are documented only in the raw schema, and the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Retrieve') plus resource ('Wave result artifact') with scope qualifiers ('one', 'principal-scoped', 'bounded'). It is distinguishable from get_job_status/get_job_result by the 'artifact' noun, though it never names those siblings to make the distinction explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all: no statement of prerequisites, no mention of when to prefer this over get_job_result or get_job_status, and no note on what state a job must be in for an artifact to exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_resultGet RQM job resultBRead-onlyIdempotentInspect
Read one completed and already-settled product result.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds the meaningful constraint that only completed/settled jobs yield a result, implying failure otherwise, but it omits error behavior, retry semantics, or result freshness. With annotations carrying most of the load, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It is arguably too terse for a tool with four schema variants, but every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover safety. What is missing is differentiation from the close siblings get_job_status and get_job_artifact and any hint about the versioned reference variants in the schema, which matters for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the anyOf branches document product/schema_version via const and enum values, so the schema does the heavy lifting. The description adds no detail about job_id format, the product variant selection, or the versioned reference shapes the schema offers. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (one product result) and qualifies scope with 'completed and already-settled', which distinguishes it in spirit from get_job_status. However, it never names get_job_artifact or get_job_status, so an agent must infer the boundary between 'result', 'status', and 'artifact' on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'completed and already-settled' implicitly defines a precondition for use, which is better than nothing. But there is no explicit when-to-use versus get_job_status (check progress) or get_job_artifact (retrieve files), nor any statement about what happens if the job is not yet settled.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_statusGet RQM job statusBRead-onlyIdempotentInspect
Read one authorized Wave or Studio managed-simulator job.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive, and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds only the word 'authorized', hinting at permission scoping, but says nothing about what state is returned or any access constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is arguably too terse given the routing ambiguity, but nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema and full annotations exist, so return values and safety need not be re-explained. Still, against a sibling set containing cancel_job, get_job_result, and get_job_artifact, the definition omits the distinctions an agent needs, leaving a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with the job reference variants (job_id, product enum, schema_version) fully typed in the input schema, so the schema does the heavy lifting. The description's 'one ... job' adds only a marginal hint that a single job is addressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('one ... job') and scopes it to Wave/Studio managed-simulator jobs, which is useful domain narrowing. However it does not distinguish itself from close siblings such as get_job_result and get_job_artifact, so an agent cannot tell from the text alone why it would pick this one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus cancel_job, get_job_result, or get_job_artifact. The description only names the object it reads, leaving all routing decisions to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_account_jobsFind this agent's purchasesARead-onlyInspect
List up to 100 recent jobs or recover the original purchase using its idempotency key.
| Name | Required | Description | Default |
|---|---|---|---|
| idempotency_key | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false; the description adds the 100-item cap and the idempotent recovery behavior, which are genuinely useful behavioral details. There is no contradiction between the description and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, action-first sentence with no wasted words. It efficiently conveys both usage modes and the result cap, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-optional-parameter read-only tool, the description covers the core behaviors well. However, there is no output schema, and the description does not indicate the shape of returned data, what 'recent' means, or behavior when the idempotency key is not found.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only defines idempotency_key as a string, and schema description coverage is 0%. The description adds meaning by connecting the key to recovering the original purchase, but it does not explicitly state that omitting the key triggers the list mode or explain where the key comes from.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List'/'recover') and a resource ('jobs'/'purchase'), and adds a clear limit of 100 recent jobs. However, it does not explicitly distinguish this tool from the similar sibling list_buyer_jobs, and the title's 'purchases' vs description's 'jobs' introduces slight terminology ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The two branches—list recent jobs versus recover the original purchase using the idempotency key—imply when the optional parameter should be used. However, there are no explicit conditions, exclusions, or comparison to alternatives, so an agent cannot tell when this tool should be preferred over list_buyer_jobs or get_account_job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_buyer_jobsList RQM buyer jobsCRead-onlyIdempotentInspect
Return complete agent discovery contracts and readiness for the exact 20/20/20 RQM specialist work portfolio.
| Name | Required | Description | Default |
|---|---|---|---|
| product | No | ||
| surface | No | buyer_job | |
| schema_version | No | rqm.jobs.agent-catalog-request.v1 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world behavior, so the safety profile is covered. The description adds no further behavioral context — no pagination, filtering, or result-size behavior — and instead makes an opaque completeness claim ('exact 20/20/20') that the agent cannot verify or act on.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, reasonably short sentence with no filler padding, which is appropriate in size. However, it front-loads cryptic jargon ('20/20/20 RQM specialist work portfolio') rather than the plain purpose, so the sentence does not fully earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but with three undocumented parameters at 0% coverage and an ambiguous description of what is actually returned, the definition is not complete enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions product, surface, or schema_version. The two enums (product, surface) and the fixed schema_version const are left entirely unexplained, so the description does nothing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name and title say 'list buyer jobs', but the description instead promises 'complete agent discovery contracts and readiness for the exact 20/20/20 RQM specialist work portfolio' — a phrase that never names the resource being listed and reads as internal jargon. An agent cannot confidently distinguish this from siblings like search_buyer_jobs or list_capabilities based on this text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this rather than search_buyer_jobs, list_capabilities, or run_buyer_job, and no prerequisites or exclusions. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_capabilitiesList RQM job capabilitiesBRead-onlyIdempotentInspect
Describe the preserved v0/v1 surface or the dynamic v2/v3 Wave catalog.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety behavior is covered. The description adds one genuine trait beyond the annotations: that v0/v1 output is 'preserved' (static) while the v2/v3 Wave catalog is 'dynamic'. That distinction is useful but thin relative to the tool's apparent complexity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words; it immediately states the two output modes. The compression comes at some cost to clarity, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required. Yet the input schema offers six const-variant branches for schema_version/product, and the description only vaguely gestures at 'v0/v1' versus 'v2/v3', leaving an agent uncertain which schema_version value to supply. Adequate but incomplete for the schema's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero counted parameters, the baseline is 4, and schema description coverage is reported at 100%. The description loosely maps its version language (v0/v1 vs v2/v3) onto schema_version variants, adding mild meaning, though it does not name the actual const strings the schema requires.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys that the tool surfaces a capability catalog and distinguishes two modes ('preserved v0/v1 surface' vs 'dynamic v2/v3 Wave catalog'), which hints at what it returns. However, the phrasing is jargon-heavy and never plainly states it lists all available job capabilities, nor does it differentiate itself from the sibling get_capability (singular). Purpose is inferable but not crisp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance is given. The two version modes are named but the description never tells the agent which to choose or when to prefer this over get_capability. Usage must be inferred entirely from the name and title.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quote_jobQuote an RQM jobCRead-onlyIdempotentInspect
Return a preserved v0 placeholder or a durable v1 Account Core USD quote.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false, so the safety profile is covered. The description does add one non-obvious behavioral fact: v0 requests yield a non-durable placeholder while v1 yields a durable Account Core USD quote, which is useful context beyond the annotations, though it does not explain what 'preserved placeholder' means operationally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler, and the v0/v1 contrast is placed early. It is efficient, though the jargon ('preserved v0 placeholder', 'Account Core USD quote') is dense for such a short statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The input schema offers five complex anyOf branches with const schema versions and enums, and an output schema exists, yet the description does not explain which branch applies, whether quoting creates a job, or what the placeholder means. For a tool this structurally complex, the description is far too thin to let an agent invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage reported at 100%, the schema carries the field-level burden. The description contributes only an oblique hint tying 'v0/v1' durability to the schema_version branches, and never clarifies product, capability_id, request, or idempotency_key. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says it returns a quote, and the title already establishes 'Quote an RQM job', but the actual action is never stated in plain language. The v0-placeholder vs v1-durable-quote distinction is a return-value nuance rather than a crisp statement of purpose, leaving the verb+resource largely inferred from the name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to quote versus when to run a job with the sibling run_buyer_job, submit_quantum_job, or submit_wave_job, nor any prerequisite or sequencing advice (e.g., quote-before-submit). The agent must guess how this fits into the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_analyze_force_disturbance_response_v1Analyze Force Disturbance ResponseCIdempotentInspect
Problem: Analyze these bounded force and disturbance inputs against the supplied response model and criteria. Input: JSON with mass kg, damping n s per m, stiffness n per m, time step seconds.... Result: modeled response, metrics and rule results. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false. The description adds useful operational limits (65536 request bytes, 5 s execution) and scopes evidence to 'software/model evidence only', which goes beyond the annotations. It does not disclose permissions or side effects of the non-read-only operation, so it is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The pseudo-structured 'Problem/Input/Result/Limits' format is compact and front-loads the problem statement, but fragments like 'time step seconds....' read as truncated. It is reasonably sized yet not cleanly prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The description covers input gist and limits, but with 0% schema coverage and nested additionalProperties, the agent lacks enough parameter detail to construct a valid request confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema does not explain the four required parameters. The description mentions a nested 'request' containing mass, damping, stiffness, and time step, which gives some domain meaning but leaves schema_version, idempotency_key, and max_total_price (including the odd const price of '0.010000') unexplained. It partially compensates but far from fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description's 'Problem' line states the tool analyzes bounded force and disturbance inputs against a supplied response model and criteria, which names a specific analytical operation. However, it is cluttered with 'Input/Result/Limits' framing rather than a clean verb+resource statement, and it does not differentiate itself from closely named siblings such as robotics_optimize_disturbance_rejection_v1 or robotics_compare_controller_responses_v1.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no named alternative among the many robotics_* siblings. The description implies the tool is for disturbance-response analysis but leaves the agent to guess whether this supersedes optimization or comparison tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_assess_calibration_quality_v1Calibration quality assessmentAIdempotentInspect
Problem: Assess this bounded calibration dataset against the supplied residual, coverage, and consistency rules. Input: JSON with length unit, observations, rules. Result: pass, fail, or indeterminate verdict, residual metrics, per-rule evaluations. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false), the description adds concrete constraints: 65536 request bytes, 5 s execution, and software/model evidence only. It does not explain the billing/job-creation implied by max_total_price and non-read-only behavior, which is a remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits framing is front-loaded, scannable, and free of filler. Every clause carries distinct information about scope, contract, output, and bounds.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested-object tool the description covers problem, input shape, result, and hard limits, and an output schema already exists so return values need not be spelled out. It remains silent on the cost/idempotency parameters, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 required params. The description documents the nested request payload (length unit, observations, rules) but says nothing about idempotency_key, schema_version, or max_total_price, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: assess a bounded calibration dataset against residual, coverage, and consistency rules. Among the many robotics_* siblings it is clearly the calibration-quality one. It stops short of naming which sibling to use instead, so full differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the input contract ('JSON with length unit, observations, rules') and the 'software/model evidence only' scoping line, but there is no explicit when-to-use / when-not-to-use guidance and no routing to alternatives such as the validation or estimation siblings. Adequate but thin.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_check_collision_constraints_v1Check Collision ConstraintsAIdempotentInspect
Problem: Check this bounded geometry and trajectory against the supplied collision and clearance constraints. Input: JSON with trajectory, obstacles, robot radius, minimum clearance. Result: clearance, collision candidates and constraint results. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses concrete limits (65536 request bytes, 5 s execution) and states it is software/model evidence only, which is real behavioral context. Annotations already cover the mutation/idempotency profile, and the description adds the evidence-scope caveat and size/time bounds.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded Problem/Input/Result/Limits structure, each line earning its place with no filler. Templates risk being terse, but here it is dense and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, inputs, result shape, and limits, and an output schema exists so return details need not be explained. But it omits usage routing among numerous robotics siblings and the meaning of the envelope params like max_total_price.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the schema is essentially an opaque request blob (additionalProperties of arbitrary JSON) plus fixed envelope fields. The description enumerates the meaningful request contents (trajectory, obstacles, robot radius, minimum clearance), which is more than the schema gives, so baseline-3-plus is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource objects (bounded geometry and trajectory against collision/clearance constraints). It is distinguishable from siblings like robotics_optimize_trajectory_to_constraints_v1, though it does not explicitly name that sibling or state the check-vs-optimize distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, when-not-to-use, or alternative-tool guidance. With many close robotics siblings (validate limits, optimize to constraints), an agent gets nothing to route by.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_compare_controller_responses_v1Compare Controller ResponsesBIdempotentInspect
Problem: Compare these bounded controller responses on identical traces using the supplied objective rules. Input: JSON with baseline, candidate, rules. Result: metric values, threshold verdicts, difference series. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (idempotentHint, destructiveHint=false), the description discloses concrete behavioral constraints: software/model evidence only, a 65536 request-byte cap, and a 5 s execution limit. These limits are genuinely useful and not derivable from the annotations, though auth or failure behavior is not covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Labeled structure (Problem/Input/Result/Limits) front-loads the purpose and constraints, and every clause carries information. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested-object, zero-schema-coverage tool, the description covers the problem, input payload, result shape, and hard limits, and an output schema exists to cover return details. The only gap is top-level parameter semantics, already scored separately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden. It names inner request fields (baseline, candidate, rules) but omits the actual required top-level parameters (schema_version, idempotency_key, max_total_price), leaving the caller's parameter picture incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (compare) and resource (bounded controller responses on identical traces using objective rules), so the operation is unambiguous. It does not, however, distinguish itself from closely named siblings like robotics_compare_state_estimators_v1 or robotics_compare_digital_twin_observations_v1.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Problem:' framing implies the use case but gives no explicit when-to-use, when-not-to-use, or alternatives among the many compare_* siblings. Nothing routes an agent between this tool and the other comparison tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_compare_digital_twin_observations_v1Compare Digital Twin ObservationsBIdempotentInspect
Problem: Compare digital-twin output with mapped observations against caller tolerances. Input: JSON with predicted, observed, declared mapping, rules. Result: typed verdict, measured metrics, candidate only when verified. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description usefully adds limits (65536 request bytes, 5 s execution), the 'software/model evidence only' scope, and 'candidate only when verified' outcome semantics. However, it never explains why a 'compare' operation is non-read-only (the paid-job/max_total_price nature implied by the schema), which is the most important behavioral fact an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits structure is front-loaded and dense with no filler. Every clause carries information, though the telegraphic style sacrifices some readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be described. The description covers purpose, limits, and result shape, but given a deeply nested schema with 4 required envelope parameters at 0% coverage, the omission of parameter semantics and the paid-job/submission mechanics leaves meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 4 required parameters, so the description must carry the load. It only loosely names request-object contents ('predicted, observed, declared mapping, rules') and never explains the required envelope fields schema_version, idempotency_key, or particularly max_total_price, whose const value and role are opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (compare) and resource scope (digital-twin output vs. mapped observations against caller tolerances), which is more concrete than a generic 'compare' tool. It does not, however, differentiate itself from near-siblings like robotics_compare_planned_observed_trajectories_v1 or robotics_compare_state_estimators_v1, so an agent must still guess which compare tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Problem:' framing gives context on what the tool is for, but there is no explicit when-to-use, when-not-to-use, or alternative routing among the many robotics_compare_* siblings. Nothing tells the agent when this is preferable to the trajectory or estimator comparison tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_compare_planned_observed_trajectories_v1Compare Planned Observed TrajectoriesAIdempotentInspect
Problem: Compare planned and observed trajectories against caller tolerances. Input: JSON with planned, observed, rules. Result: typed verdict, measured metrics, candidate only when verified. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses useful behavioral limits: software/model evidence only, 65536 request bytes, and 5 s execution. It also describes the result as a typed verdict with measured metrics, and notes that a candidate is returned only when verified. It does not explain the readOnlyHint=false implication, but it adds substantial operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded using labeled fragments: Problem, Input, Result, Limits. Every sentence carries information, and there is no redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need not be fully explained, and the description covers result type at a high level. However, the input schema is opaque and the description does not clarify the required top-level request fields or how to choose this tool over sibling comparison tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for 4 required parameters. The description only names planned, observed, and rules as inputs, leaving schema_version, idempotency_key, max_total_price, and the exact structure of the request object unexplained. It adds some meaning but does not compensate for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: comparing planned and observed trajectories against caller tolerances. This is naturally distinct from sibling comparison tools such as compare_controller_responses and compare_state_estimators, so an agent can identify the tool without opening its schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The problem statement implies the use case, but there is no explicit when-to-use guidance, no exclusion conditions, and no mention of alternatives among the many sibling comparison tools. An agent must infer selection from the tool name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_compare_state_estimators_v1Compare State EstimatorsCIdempotentInspect
Problem: Compare these bounded estimator outputs against the supplied reference and objective rules. Input: JSON with reference, estimators, rules. Result: per-estimator metrics, rule verdicts, deterministic selected estimator or.... Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false, so the safety posture is structured. The description adds genuinely useful constraints: software/model evidence only, 65536 request bytes, 5 s execution, and a deterministic selected estimator. These are extra behavioral facts beyond the annotations, which lifts it above baseline, though mutation and reversibility are not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loads the core operation, but the telegraphic 'Problem/Input/Result/Limits' labels and the trailing 'or....' make it read as an unfinished template rather than polished prose. Density is acceptable but the structure is awkward.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain return values, and it correctly sketches the result shape. However, for a tool with four required, undocumented parameters and a nested opaque request object, plus many sibling robotics tools, the definition is thin on how to construct inputs and when to prefer it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four required parameters. The description mentions 'JSON with reference, estimators, rules' which loosely maps to the opaque 'request' object, but it does not explain schema_version, idempotency_key, or max_total_price, nor the required structure inside request. With 0% coverage, the description is expected to compensate and largely does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb+resource: compare bounded estimator outputs against a supplied reference and objective rules, producing metrics and a selected estimator. This is more specific than the title, but it uses a 'Problem/Input/Result/Limits' template that reads like a schema stub rather than a clear purpose statement, and it does not differentiate itself from sibling robotics comparison tools like robotics_compare_controller_responses_v1 or robotics_estimate_fused_state_v1.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative guidance is provided. With many sibling robotics_* comparison/estimation tools, the agent has no signal for choosing this tool over the others, and the 'Limits' line only constrains evidence type and size.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_compute_actuator_alignment_correction_v1Compute Actuator Alignment CorrectionBIdempotentInspect
Problem: Compute a bounded actuator alignment correction from these observations and caller acceptance limits. Input: JSON with observations, model, rules. Result: correction parameters, pre/post residuals, applicability verdict. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=false. The description usefully adds the 65536-byte request cap and 5 s execution limit, which are not in structured data. However, it omits that this is a priced job (max_total_price pinned to 0.010000), that an idempotency key is required, and any cost/funding implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly labelled clauses, front-loaded with the problem statement and no filler. The telegraphic 'Input: JSON with...' style is efficient, though it sacrifices some readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return details need not be spelled out, and the limits are a nice addition. Still, for a non-read-only, priced, idempotency-keyed job submission with a fully opaque 0%-documented parameter set, the description leaves the agent without enough to construct a valid request.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 required parameters. The description names three conceptual inputs (observations, model, rules) that plausibly live inside the opaque 'request' object, but it says nothing about schema_version, idempotency_key, or max_total_price. It adds only partial meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Compute a bounded actuator alignment correction') and identifies the input class (observations, model, rules) and the output (correction parameters, pre/post residuals, applicability verdict). It reads distinctly from most siblings, though it doesn't name a near alternative such as tune_controller_to_objectives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The Problem/Input/Result/Limits frame implies when the tool applies, and 'Software/model evidence only' is a mild scope exclusion. But there is no explicit when-to-use versus sibling tools like assess_calibration_quality or tune_controller_to_objectives, and no prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_diagnose_trajectory_instability_v1Trajectory instability diagnosisAIdempotentInspect
Problem: Diagnose this bounded trajectory for the supplied oscillation, divergence, and stability indicators. Input: JSON with frame id, length unit, samples, rules. Result: pass, fail, or unsupported verdict, indicator series, violation windows. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, destructiveHint=false, idempotentHint=true, leaving the nature of the operation unexplained; the description fills that in with 'software/model evidence only', a 65536-byte request cap, and a 5 s execution limit. It also previews the verdict space (pass/fail/unsupported), adding behavior beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits layout is front-loaded and every clause carries information, with no filler sentences. It is telegraphic in places, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex nested-request tool with an output schema (so return values need not be detailed), the description still covers the input shape, verdict outcomes, and hard limits. The remaining gap is the absence of guidance against the many similar robotics siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the schema only exposes an opaque request bag plus schema_version, idempotency_key and max_total_price, so the description must carry the load. 'JSON with frame id, length unit, samples, rules' names the meaningful inner payload fields, partly compensating, but gives no types, formats, or constraints for them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (diagnose) and resource (bounded trajectory for oscillation, divergence and stability indicators), which clearly separates it from vague siblings. It stops short of explicitly distinguishing itself from close neighbors like robotics_validate_closed_loop_stability_v1 or robotics_validate_trajectory_timing_v1, so it is clear but not fully sibling-differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the many overlapping robotics validation/diagnosis siblings. The phrase 'this bounded trajectory' implies a precondition but names no exclusions or alternative tool. An agent gets no routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_estimate_fused_state_v1Estimate Fused StateBIdempotentInspect
Problem: Fuse these bounded synchronized sensor observations under the supplied diagonal uncertainty model and rules. Input: JSON with observations, model, rules. Result: fused state, diagonal covariance, sensor residuals and verdict. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, idempotent=true, destructive=false, and closed-world, so the description need not restate safety. It usefully adds hard limits (65536 request bytes, 5 s execution) and the 'software/model evidence only' scope, but omits that this is a priced buyer job (schema carries max_total_price and idempotency_key), which is meaningful behavioral context an agent needing cost awareness would want.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits skeleton is tightly front-loaded and every clause carries information. Slightly dense and telegraphic, but no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be spelled out, and the description still summarizes them well. However, for a paid, idempotent job-submission tool with a fully undocumented nested request schema, the definition omits pricing/cost behavior and job lifecycle details that materially affect correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 required params, and the schema is an opaque nested anyOf object, so the description carries real burden. It adds meaningful content by naming the payload parts ('observations, model, rules'), partially compensating for the undocumented 'request' object, but says nothing about schema_version, idempotency_key, or max_total_price.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Fuse ... synchronized sensor observations') and scopes it to a bounded observation set under 'the supplied diagonal uncertainty model and rules.' It is clear what the tool does, but it does not distinguish itself from close siblings like robotics_compare_state_estimators or robotics_validate_closed_loop_stability, which also consume sensor/state inputs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this over alternatives, no prerequisites, and no exclusions. The 'Problem:' framing implies the scenario but never states the conditions that select this estimator over a competing one among the ~20 robotics siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_frame_convention_validation_v1Frame and convention validationBIdempotentInspect
Problem: Validate this bounded frame graph and report unit, axis, handedness, rotation, and transform-direction violations. Input: JSON with frames, declared length unit, declared transform direction, declared.... Result: pass, fail, or unsupported verdict, exact findings, normalized evidence. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare this is not read-only but not destructive and is idempotent, and the description usefully adds operational limits (software/model evidence only, 65536 request bytes, 5 s execution) plus the verdict space (pass/fail/unsupported). However, it never mentions that each call carries a monetary cost (max_total_price is required) or that the readOnlyHint=false reflects a billed job submission, which is the most important behavioral fact for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits labeling gives a clear, front-loaded structure, and the limits are stated economically. Against that, the Input sentence is truncated with an ellipsis and literal 'declared....', which reads as unfinished text rather than deliberate brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return format need not be spelled out, and the description does cover evidence scope and execution limits. It remains incomplete for a billed, idempotency-keyed job tool: cost, idempotency semantics, and the shape of the request body are all unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and all four parameters are required, yet the description only loosely gestures at payload structure ('JSON with frames, declared length unit, declared transform direction, declared....'), trailing off mid-sentence. Required fields like idempotency_key, schema_version, and max_total_price are never explained, and the nested object is not described, so the description does not close the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (validate a bounded frame graph) and enumerates the exact violation classes it reports: unit, axis, handedness, rotation, and transform-direction. This readily separates it from siblings like robotics_normalize_sensor_frames_v1 (which normalizes) and robotics_assess_calibration_quality_v1 (which assesses quality).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no routing against the many sibling robotics_* validation tools. The 'Problem:' framing implies a validation context but never states the conditions that select this tool over robotics_validate_kinematic_dynamic_limits_v1 or robotics_normalize_sensor_frames_v1.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_normalize_sensor_frames_v1Sensor frame normalizationAIdempotentInspect
Problem: Convert these bounded sensor observations into the supplied target frame and unit convention. Input: JSON with frame graph, target frame id, observations. Result: normalized or unsupported verdict, normalized observations, transform paths. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal idempotency and non-destructiveness, and the description adds concrete operational context beyond them: the 65536-byte request limit, 5-second execution budget, software/model-evidence-only scope, and the normalized/unsupported verdict outcome. It omits that this is a paid buyer-job call (implied by max_total_price) and any auth/pricing behavior, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits labeling front-loads the purpose and constraints with zero filler; every clause carries information an agent needs. Well sized for a tool with a nested opaque request payload.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return contents need not be spelled out, yet the description still sketches the result. Missing pieces for a nested, paid, job-creating tool are the workflow context (that it submits a buyer job) and semantics of the required wrapper parameters. Adequate but with visible gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully characterizes the request content (frame graph, target frame id, observations), which the schema leaves fully opaque, but says nothing about the three wrapper fields (schema_version, idempotency_key, max_total_price) or the const/pattern constraints, so param semantics remain only partly covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (convert/normalize) and resource (bounded sensor observations into the supplied target frame and unit convention), which an agent can act on. It does not explicitly differentiate itself from the closely named sibling robotics_frame_convention_validation_v1, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Problem:' framing implies when the tool applies, but there is no explicit when-to-use guidance, no exclusions, and no mention of the adjacent alternatives (frame_convention_validation, estimate_fused_state, etc.). Nothing steers the agent to this tool over its many robotics siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_optimize_disturbance_rejection_v1Optimize Disturbance RejectionBIdempotentInspect
Problem: Tune supported control parameters against this bounded disturbance model and supplied response objectives. Input: JSON with mass, damping, disturbance, time step seconds, baseline kp, baseline.... Result: candidate parameters, comparative metrics and constraint results. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety/idempotency profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false), and the description adds real context beyond them: a 65536-byte request cap, a 5 s execution limit, and a 'software/model evidence only' disclaimer. It does not contradict annotations and usefully sets expectations about cost and validity of results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits framing is well-structured and front-loaded with the core purpose. It is compact, though the dangling 'baseline....' ellipsis is sloppy and leaves a partially stated input list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so explaining return values is not required, and the description does summarize the result shape anyway. However, for a nested free-form request object at 0% schema coverage, the input contract described is too thin, and key required fields like idempotency_key and the price cap are left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 4 required parameters, so the description must compensate. It does partially—naming request fields like mass, damping, disturbance, time step seconds, and baseline kp—but it trails off with an ellipsis ('baseline....') and says nothing about idempotency_key or the max_total_price semantics, so the compensation is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Tune supported control parameters') and scopes it to a bounded disturbance model against supplied response objectives. This is clear enough to act on, but it never distinguishes itself from the closely related sibling robotics_tune_controller_to_objectives_v1, leaving the agent to guess which tuning tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or alternative-routing guidance. With multiple tuning/validation siblings (tune_controller_to_objectives, validate_closed_loop_stability, compare_controller_responses), the description gives no signal for choosing among them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_optimize_trajectory_to_constraints_v1Optimize Trajectory To ConstraintsAIdempotentInspect
Problem: Optimize a bounded trajectory only when speed and duration constraints pass. Input: JSON with trajectory, maximum speed, maximum duration s. Result: typed verdict, measured metrics, candidate only when verified. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotent=true, destructive=false, closed-world, and that it is a non-read operation. The description adds genuinely new behavioral context: request size limit (65536 bytes), 5 s execution bound, 'software/model evidence only', and that a candidate is returned only when verified. It still does not explain what is mutated or authorization needs, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The terse Problem/Input/Result/Limits structure is well front-loaded and dense with no filler. Every clause carries information, though the telegraphic style is slightly terse for a complex optimization tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description covers input content, execution limits, and the conditional-verdict behavior. It is nearly complete for a mutation tool, missing only explicit prerequisite/permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 required params, so the description must compensate. It usefully characterizes the payload (trajectory, maximum speed, maximum duration) but does not explain idempotency_key, max_total_price, or schema_version semantics, so it only partially fills the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Optimize) and resource (a bounded trajectory) and scopes it with the gating condition that speed and duration constraints must pass. This is clearly distinguishable from sibling validation tools by its optimization outcome. However, it does not name or contrast any specific sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a precondition (only when speed and duration constraints pass) but gives no explicit when-to-use guidance versus alternatives. With many related siblings (validate_trajectory_timing, reconstruct_smoothed_trajectory, validate_kinematic_dynamic_limits), the agent gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_reconstruct_smoothed_trajectory_v1Reconstruct Smoothed TrajectoryBIdempotentInspect
Problem: Reconstruct this bounded trajectory under the supplied smoothing model and residual rule. Input: JSON with samples, model, rules. Result: smoothed trajectory, residual RMSE, caller-rule verdict. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give the safety profile (non-read-only, idempotent, non-destructive), so the bar is lower, and the description still adds concrete limits beyond them: software/model evidence only, a 65536-byte request cap, and a 5-second execution bound. It does not, however, explain the job-submission/lifecycle behavior implied by idempotency_key and max_total_price.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Uses a terse Problem/Input/Result/Limits structure with the core action front-loaded and no filler. Every clause carries information, though the labeled fragments are dense rather than flowing prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the description covers problem, input, result and limits. However, for a job-submission tool with a nested request object and four required params, it omits the async/paid-job context the schema implies, leaving it only minimally sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden. It partially names request contents (samples, model, rules) but leaves the other required top-level parameters (schema_version, idempotency_key, max_total_price) unexplained, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (reconstruct) and resource (bounded trajectory) plus the governing model and residual rule, which distinguishes it from siblings like robotics_optimize_trajectory_to_constraints_v1 or robotics_compare_planned_observed_trajectories_v1. It is clear but does not explicitly name those siblings to route the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Problem' framing gives context and the 'Input' line names required data (samples, model, rules), but there is no when-to-use/when-not guidance and no mention of alternatives among the many trajectory-related siblings. The agent must infer the appropriate scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_replay_workcell_trace_v1Replay Workcell TraceAIdempotentInspect
Problem: Replay this bounded workcell trace against the supplied event assertions. Input: JSON with model id, events, assertions. Result: timeline, assertion results and trace digest. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is not read-only, is idempotent, not destructive, and closed-world. The description adds valuable operational context not in the annotations: software/model evidence only, a 65536-request-byte limit, and a 5-second execution limit. This exceeds the annotation-covered baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very compact and front-loaded, using a Problem/Input/Result/Limits structure that lets an agent extract the purpose, inputs, outputs, and constraints in one pass with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested request object with no schema descriptions and an output schema present, the description supplies the core input/output gist, constraints, and evidence scope. It omits parameter-level detail, which leaves a gap against the 0% schema coverage, but the presence of an output schema reduces the need to describe results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description only loosely mentions 'JSON with model id, events, assertions' without mapping these to the actual required schema fields (request, schema_version, idempotency_key, max_total_price). With low coverage, the description should compensate but does not explain parameter meaning or format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Replay') and resource ('bounded workcell trace') with the goal of testing against supplied event assertions. Distinguishes itself from siblings like robotics_validate_* and robotics_compare_* by being a trace replay against assertions, though it does not explicitly name which sibling to prefer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'Problem:' implies the task context, but there is no explicit when-to-use, when-not-to-use, or named alternative among the many robotics_* siblings. Usage is only inferable from the described input/output.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_tune_controller_to_objectives_v1Tune Controller To ObjectivesCIdempotentInspect
Problem: Tune a supported controller only within supplied bounds and safety limits. Input: JSON with supported controller, current gain, fixture error, target error.... Result: typed verdict, measured metrics, candidate only when verified. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is largely covered. The description adds useful context beyond that: the 65536-byte request cap, 5 s execution time, and 'candidate only when verified' conditionality. These are genuine behavioral details, though several (idempotency, non-destructiveness) are already in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The label-based format (Problem/Input/Result/Limits) is telegraphic and fragmented rather than front-loaded prose. The 'Input:' line trails off with '...' and mixes object contents with parameter names. Reads as compressed notes, not a coherent definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested objects, four required params, no schema descriptions, and 65536-byte/5 s limits, the description is under-specified. Having an output schema mitigates the need to explain return values, but the input contract remains opaque and no usage guidance anchors it against ~20 siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and four required parameters (schema_version, request, idempotency_key, max_total_price) have no descriptions anywhere. The description only lists generic input concepts ('supported controller, current gain, fixture error, target error') that do not match the actual parameter names, leaving the request payload structure unexplained. Does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (tune) and resource (controller) constrained to objectives, distinguishing it somewhat from siblings. However, the 'supported controller' phrasing is vague and the sentence is fragmented into Problem/Input/Result/Limits labels rather than a clean purpose statement. Enough to identify the tool but not sharp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never states when to use this tool versus the many robotics_* siblings like robotics_optimize_disturbance_rejection_v1 or robotics_compare_controller_responses_v1. The 'Limits' line gives scope constraints (software/model evidence only) but no positive usage guidance or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_validate_closed_loop_stability_v1Validate Closed Loop StabilityBIdempotentInspect
Problem: Accept or reject this bounded closed-loop model under the supplied stability criteria. Input: JSON with discrete state matrix, criteria. Result: eigenvalues, spectral radius and rule results. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false. The description adds the limits: 65536 request bytes and 5 s execution, which are useful constraints not covered by annotations. It also mentions 'Software/model evidence only,' indicating the scope. However, it doesn't explain the implications of readOnlyHint=false (likely a job submission that consumes resources) or the idempotency behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, using a structured format (Problem, Input, Result, Limits) that front-loads the key information. Every sentence adds value without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (4 required parameters, nested objects, no parameter descriptions, output schema exists), the description is incomplete. It doesn't cover the parameters, the meaning of readOnlyHint=false, or how to construct the request. While the output schema exists, the input parameters remain unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'JSON with discrete state matrix, criteria' but this doesn't map to the actual parameters (request, schema_version, idempotency_key, max_total_price). The description fails to explain these parameters, especially max_total_price which has a const value of 0.010000 and is likely important for cost control.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific purpose: accept or reject a bounded closed-loop model under supplied stability criteria. It's clear what the tool does, though it doesn't explicitly differentiate from siblings like robotics_diagnose_trajectory_instability_v1 or robotics_validate_kinematic_dynamic_limits_v1.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. There's no mention of prerequisites or scenarios where other validation tools are more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_validate_kinematic_dynamic_limits_v1Validate Kinematic Dynamic LimitsAIdempotentInspect
Problem: Accept or reject a bounded trajectory against caller kinematic limits. Input: JSON with trajectory, maximum speed, maximum acceleration. Result: typed verdict, measured metrics, candidate only when verified. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and readOnlyHint=false, and the description reinforces this with the bounded 'software/model evidence only' scope, a 65536-byte request cap, a 5 s execution window, and the 'candidate only when verified' verdict semantics. It does, however, leave undisclosed that this schema is a paid buyer-job submission (required max_total_price), which is a notable behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tight, front-loaded Problem/Input/Result/Limits framing with no filler; every clause conveys distinct operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be restated, and the description still covers problem, input, result, and hard limits for a nested-object tool. The main omission is that this is a paid job submission requiring max_total_price.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It usefully maps the request payload (trajectory, maximum speed, maximum acceleration) but leaves the required top-level parameters schema_version, idempotency_key, and especially the fixed max_total_price completely unexplained, so compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: accept or reject a bounded trajectory against caller-supplied kinematic limits, plus the expected outputs (typed verdict, measured metrics). It is clearly a validation tool, though it does not explicitly differentiate itself from nearby siblings like robotics_validate_trajectory_timing_v1 or robotics_check_collision_constraints_v1.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Problem' line implies when the tool applies (checking a bounded trajectory against kinematic limits), but there is no explicit when-to-use guidance, no prerequisites, and no reference to alternatives such as the timing or collision sibling validators.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
robotics_validate_trajectory_timing_v1Trajectory and sensor timing validationBIdempotentInspect
Problem: Validate timestamps, ordering, rates, gaps, and synchronization in this bounded trajectory or sensor trace. Input: JSON with clock id, samples, rules. Result: pass or fail verdict, normalized timeline, exact timing defects. Limits: Software/model evidence only; 65536 request bytes; 5 s execution.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotent=true and destructive=false, and the description adds genuine operational context beyond them: software/model-only evidence, a 65536-byte request cap, a 5s execution limit, and the fact that the result is a pass/fail verdict plus normalized timeline. It stops short of explaining the pricing/idempotency-key obligations implied by readOnlyHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Problem/Input/Result/Limits structure is tightly front-loaded and wastes no words; limits and outcomes are surfaced early. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not strictly required, and the description usefully covers the payload shape and execution limits. But for a tool with nested objects, a 0%-coverage schema, and three undocumented buyer-job wrapper parameters, it leaves the agent under-informed about required non-obvious fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 required top-level parameters. The description characterizes the request payload (clock id, samples, rules), which adds real meaning for the nested 'request' object, but schema_version, idempotency_key, and max_total_price (a fixed '0.010000' const) are left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (validate) on a specific resource (timestamps, ordering, rates, gaps, synchronization in a trajectory/sensor trace), which distinguishes it from siblings like robotics_validate_kinematic_dynamic_limits_v1 and robotics_frame_convention_validation_v1. It is clear what the tool checks, though the sibling differentiation is implicit rather than named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It scopes usage to 'this bounded trajectory or sensor trace' but offers no when-to-use/when-not guidance and never names an alternative among the many robotics_* siblings. The agent must infer selection from the purpose statement alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_account_jobRun a job using prepaid RQM creditsADestructiveIdempotentInspect
Submit one catalog job within the account owner's saved spending policy. Reuse idempotency_key on every retry. This spends prepaid credits; it never signs x402 or funds the account.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| capability_id | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructive/idempotent/not-read-only. The description adds concrete behavior: it spends prepaid credits, never signs x402 or funds the account, and instructs idempotency_key reuse. This goes beyond the structured annotations by specifying what the destructive action actually is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each with a clear purpose: main action, idempotency retry rule, and behavioral boundary. No filler or redundancy; the description is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 required parameters, a nested request object, and no output schema, the description is too thin. It omits parameter meaning (capability_id, request, max_total_price details), doesn't mention prerequisites or expected return behavior, and doesn't compensate for the 0% schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It only clarifies idempotency_key explicitly and hints at max_total_price via 'spending policy.' The essential parameters capability_id and request are left unexplained, leaving the agent to guess the request shape and capability semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb and resource: 'Submit one catalog job within the account owner's saved spending policy.' This distinguishes it from siblings like run_buyer_job (buyer vs account) and submit_quantum_job/submit_wave_job (catalog vs other job types). The title reinforces the prepaid-credit aspect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: account owner's catalog job, prepaid credits, saved spending policy. Also gives a retry instruction (reuse idempotency_key) and states non-behaviors (never signs x402 or funds the account). However, it does not explicitly name alternatives or say when to pick this over run_buyer_job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_buyer_jobRun any RQM buyer jobBIdempotentInspect
Quote, settle, and submit one allowlisted WaveEngine, Studio, or Robotics buyer job using x402.
| Name | Required | Description | Default |
|---|---|---|---|
| product | Yes | ||
| request | Yes | ||
| capability_id | Yes | ||
| schema_version | Yes | ||
| idempotency_key | Yes | ||
| max_total_price | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, idempotent, non-destructive, closed-world. The description adds meaningful behavior beyond that: this is a three-phase flow that actually settles payment (spends funds), which is critical context the annotations' destructiveHint=false could otherwise obscure. It still omits what happens on partial failure across quote/settle/submit and whether errors are recoverable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the three phases and then the target resource. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be described. But for a six-required-parameter, payment-settling, composite tool with an opaque request object, the definition leaves too much unexplained: no parameter meaning, no ordering between quote/settle/submit, and no routing against the many competing siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for six required parameters, and the description adds no meaning for any of them. In particular max_total_price (the spend cap), capability_id (allowlisted capability selector), idempotency_key, and the free-form request object are given no semantics in either place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific composite action set (quote, settle, submit) on a concrete resource (one allowlisted WaveEngine/Studio/Robotics buyer job) via x402. An agent can tell it is the end-to-end job runner, but the jargon 'allowlisted' and 'x402' and the overlap with sibling quote_job / submit_wave_job / submit_quantum_job are not resolved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given. The tool overlaps heavily with quote_job (quote only), submit_wave_job and submit_quantum_job (product-specific submit), yet the description never says when to prefer this composite runner over those narrower siblings, nor what prerequisites (allowlisting, funding session) must exist first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_buyer_jobsSearch RQM buyer jobsCRead-onlyIdempotentInspect
Select RQM buyer jobs from work-language requests using deterministic semantic scoring over instructions, examples, outputs, and next steps.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| product | No | ||
| surface | No | buyer_job | |
| schema_version | No | rqm.jobs.agent-search-request.v1 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds one genuine behavioral trait beyond that: retrieval is 'deterministic semantic scoring' over 'instructions, examples, outputs, and next steps.' It does not disclose ranking guarantees, tie-breaking, or how limit interacts with scoring.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted filler. It is dense but efficient; the mild cost is that compactness comes at the expense of the parameter detail an agent actually needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a 5-parameter search tool with 0% schema coverage the description leaves key inputs (surface, product) and their effect undocumented. An agent cannot determine how to scope a query from this definition alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description carries the full explanatory burden and fails to meet it. It never explains 'surface' (buyer_job/implementation_capability/research_evidence/all), 'product', 'limit', or 'schema_version' – the single most consequential parameter, 'surface', is completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Select RQM buyer jobs ... using deterministic semantic scoring.' An agent can tell it is a relevance-search tool, distinct from list_buyer_jobs by implication of scoring rather than enumeration. However, it never names the sibling it competes with, and 'from work-language requests' is jargon that blurs the actual input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not guidance. The obvious alternative, list_buyer_jobs, is never mentioned, nor is the condition (semantic/ranked retrieval vs. plain enumeration) that would select this tool. Usage is only inferable from the phrase 'semantic scoring.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
settle_buyer_jobAuto-settle a Jobs buyer 402AIdempotentInspect
Server-side Jobs-rail x402 settle for connector hosts that cannot sign. Replays the same idempotency key with PAYMENT-SIGNATURE. An RQM-owned operator signer is not external revenue.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, readOnlyHint=false, openWorldHint=false. The description goes beyond them by disclosing the mechanism: it replays the same idempotency key with PAYMENT-SIGNATURE server-side, explaining how idempotency and signing are achieved. The final clause about the operator signer is cryptically worded but adds a revenue/ownership caveat.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and no filler. The closing sentence ('An RQM-owned operator signer is not external revenue') is opaque and its purpose is unclear, a minor structural blemish on an otherwise tight definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Return values need not be described given an output schema exists, and annotations cover the safety profile. For a payment-settlement tool the description conveys what it does, when (hosts that cannot sign) and the key mechanism, though the request-variant semantics rest entirely on the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The report lists 0 parameters with 100% schema coverage, so the oneOf challenge/capability structure (idempotency_key, payment_required, capability_id) is already documented in the schema. The description references idempotency-key replay and PAYMENT-SIGNATURE, loosely reinforcing schema concepts, but does not explain the two request variants. Baseline 4 applies for the zero-parameter case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource scoped to a niche: 'Server-side Jobs-rail x402 settle'. Combined with the title 'Auto-settle a Jobs buyer 402', an agent can tell this handles payment settlement rather than job execution like run_buyer_job. The heavy acronym/jargon load (x402, RQM, PAYMENT-SIGNATURE) makes intent clear only to a domain-aware reader, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for connector hosts that cannot sign' gives a real usage condition, and 'An RQM-owned operator signer is not external revenue' hints at an exclusion. However, no sibling alternative (e.g., run_buyer_job, quote_job) is named and no explicit when-not is given, so routing is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_quantum_jobSubmit an RQM Studio jobAIdempotentInspect
Reserve an exact quote and submit one bounded managed-simulator job.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a write (readOnlyHint false), safe (destructiveHint false), idempotent, and closed-world. The description adds that a quote must be reserved and that the job is bounded and managed-simulator scoped, which is real behavioral context beyond annotations, though it omits auth, cost, and rate-limit details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence beginning with the action and state change, with no filler. It is appropriately terse given the rich schema and annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Core purpose and a key prerequisite are stated, and the output schema plus annotations carry return and safety details. However, the description does not explain the three anyOf schema variants or clarify that capability_id accepts many values beyond managed-simulator-v1, leaving an agent to rely entirely on the schema for that complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description only indirectly signals quote_id ('exact quote') and capability_id ('managed-simulator'), adding little meaning beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('submit') and resource ('job'), with scope ('one bounded managed-simulator job') and a prerequisite ('Reserve an exact quote'). This distinguishes it from the batch sibling submit_wave_job via 'one' and from run_buyer_job, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Signals an implied prerequisite—reserving a quote—which maps conceptually to the sibling quote_job, but never states when to use this tool versus run_buyer_job or submit_wave_job, nor any exclusions. Usage is inferable but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_wave_jobSubmit a WaveEngine jobBIdempotentInspect
Reserve an exact quote and submit one bounded catalog-approved Wave job.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-readonly, idempotent, non-destructive, closed-world write, so safety is covered. The description adds that a quote is reserved and the job is bounded/single, but says nothing about idempotency-key behavior, failure/retry outcomes, or what a reservation costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though its brevity borders on underspecification for a fairly complex anyOf schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema is a three-variant anyOf with version consts (v1/v2) and a constrained capability_id, and an output schema exists so returns are covered. However, the description never acknowledges these variants or the required fields, leaving the agent to discover the version/capability coupling purely from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported as 100%, so the schema is expected to carry parameter meaning and the baseline is 3. The description adds no semantics for quote_id, capability_id, schema_version variants, or idempotency_key beyond what the structured fields imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (submit) and resource (Wave job), with scoping qualifiers ('one bounded catalog-approved'). The 'Wave' qualifier implicitly separates it from submit_quantum_job, though the description never names the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Reserve an exact quote' and 'catalog-approved' imply a prerequisite workflow (a quote must exist before submission), but the description gives no explicit when-to-use, no exclusions, and never contrasts with quote_job or submit_quantum_job. Usage is only inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- Added
get_account_job - Added
list_account_jobs - Added
run_account_job
1 tool update
- Added
settle_buyer_job
34 tool updates
- First observed
cancel_job - First observed
create_funding_session - First observed
get_account_balance - First observed
get_capability - First observed
get_job_artifact - First observed
get_job_result - First observed
get_job_status - First observed
list_buyer_jobs - First observed
list_capabilities - First observed
quote_job - First observed
robotics_analyze_force_disturbance_response_v1 - First observed
robotics_assess_calibration_quality_v1 - First observed
robotics_check_collision_constraints_v1 - First observed
robotics_compare_controller_responses_v1 - First observed
robotics_compare_digital_twin_observations_v1 - First observed
robotics_compare_planned_observed_trajectories_v1 - First observed
robotics_compare_state_estimators_v1 - First observed
robotics_compute_actuator_alignment_correction_v1 - First observed
robotics_diagnose_trajectory_instability_v1 - First observed
robotics_estimate_fused_state_v1 - First observed
robotics_frame_convention_validation_v1 - First observed
robotics_normalize_sensor_frames_v1 - First observed
robotics_optimize_disturbance_rejection_v1 - First observed
robotics_optimize_trajectory_to_constraints_v1 - First observed
robotics_reconstruct_smoothed_trajectory_v1 - First observed
robotics_replay_workcell_trace_v1 - First observed
robotics_tune_controller_to_objectives_v1 - First observed
robotics_validate_closed_loop_stability_v1 - First observed
robotics_validate_kinematic_dynamic_limits_v1 - First observed
robotics_validate_trajectory_timing_v1 - First observed
run_buyer_job - First observed
search_buyer_jobs - First observed
submit_quantum_job - First observed
submit_wave_job
Related MCP Connectors
Delegate 20 typed RQM Robotics jobs for validation, diagnosis, comparison, and optimization.
Delegate 20 fail-closed RQM Studio quantum work products with fixed-request x402.
Delegate 20 typed, evidence-backed WaveEngine signal work products with fixed-request x402.
Delegate tasks to vetted human experts - research, writing, analysis, and data work.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables routing tasks to the right specialist agents, assembling the needed context from a layered catalog of agents, skills, knowledge, runbooks, topics, tools, and connectors, while remembering how a business wants work done.1013 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables autonomous agents to manage tasks in a pull-based work queue with strategic goal alignment, real-time monitoring, and cross-project choreography.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI to delegate boilerplate, drafts, tests, and refactors to free LLM providers, saving tokens and running tasks in parallel.1,024 npmMIT
- FlicenseNot gradedqualityCmaintenanceSpawn headless OpenCode workers on cheap models to delegate expensive tool-using tasks to a fleet of parallel workers with isolated git worktrees.-
Glama MCP Gateway
Add one secure layer between your agents and this server.