Uxia
Server Details
Uxia brings AI user testing, usability testing and UX research into ChatGPT. It helps product managers, UX designers, researchers and developers evaluate websites, interactive prototypes and digital product experiences using AI-simulated users. Explore potential usability problems, investigate friction in user journeys and use findings to guide product and design improvements.
- Status
- Healthy
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 22 tools
There is significant overlap among the many test-reading tools (get_test, get_test_definition, get_test_progress, list_test_blocks, list_tests), which could lead to misselection. Similarly, multiple editing tools (add_task, update_task, set_survey, set_test_audience, edit_draft_advanced) target draft modifications, though descriptions help distinguish them. Overall, some ambiguity remains.
All tool names follow a consistent verb_noun snake_case pattern, such as add_task, create_audience, get_test_progress, and list_insights. The few multi-word names like create_block_test and preview_test_launch still adhere to the same convention, making the set highly predictable.
With 22 tools, the count falls in the borderline-heavy range (16-25) for the apparent platform scope. While the complexity of managing tests, audiences, drafts, and insights justifies many tools, several read and edit operations are granular enough that the set feels somewhat bloated.
The server covers creation, reading, and updating for tests and audiences, but lacks any explicit delete operations for tests, audiences, or blocks. This notable gap means agents cannot fully manage lifecycle without resorting to advanced operations (find_operations) or leaving orphaned resources.
Available Tools
22 toolsadd_taskBIdempotentInspect
Add a website task or empty task to an existing draft. Flat fields; no configuration object. Omit wording only for an intentionally incomplete draft. With url, supply the user-established isPrototype flag. Use find_operations for prototypes or advanced settings. No launch.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| task | No | ||
| testId | Yes | ||
| blockId | Yes | ||
| position | No | ||
| scenario | No | ||
| isPrototype | No | ||
| stopCondition | No | ||
| expectedRevision | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context ('No launch', the isPrototype requirement tied to url), but does not explain mutation effects, revision handling, or what happens to existing tasks. Moderate added value over annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and compact, with the core action stated first and rules following. Sentence fragments ('No launch.') are efficient rather than wasteful, though 'Flat fields; no configuration object' borders on cryptic without explanation of what configuration would otherwise be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, with 9 parameters including three required identifiers (testId, expectedRevision, blockId) and 0% schema coverage, the description is not complete enough to reliably construct a call. It covers the differentiating fields but omits mandatory and optional semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description carries the full burden. It explains url, isPrototype, and 'wording' (task), but leaves testId, expectedRevision, blockId, position, scenario, and stopCondition entirely undocumented. Roughly two-thirds of parameters get no semantic help.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('add') and resource ('task') with clear scope ('to an existing draft'), and distinguishes itself from edit_draft_advanced by noting 'flat fields; no configuration object'. It clearly signals it is not a launch operation, separating it from launch_test. Sibling differentiation is present though not exhaustive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives contextual rules ('Omit wording only for an intentionally incomplete draft') and an explicit alternative ('Use find_operations for prototypes or advanced settings'), plus a scoping statement ('No launch'). This approaches when/when-not guidance, though it never states prerequisites like draft revision state or permissions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_audienceCreate AudienceAInspect
Create a new audience in the caller's Uxia workspace and synchronously generate AI testers for it. An audience defines targeting filters (age, gender, geography, demographics) used to recruit testers. Returns the created audience plus the freshly generated testers.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| count | No | How many AI testers to generate. Defaults to 5, which is the minimum a test can launch with — an audience with fewer approved testers is rejected at launch. | |
| roles | No | ||
| gender | Yes | ||
| ageRange | Yes | Two-element [min, max] age range (integers, 0-120). | |
| countries | No | ||
| languages | No | ||
| industries | No | ||
| isReusable | No | ||
| description | No | ||
| techLiteracy | No | ||
| customDetails | No | ||
| maritalStatus | No | ||
| nationalities | No | ||
| educationLevel | No | ||
| personalIncome | No | ||
| householdIncome | No | ||
| employmentStatus | No | ||
| languagesRequirement | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| testers | No | Freshly generated AI testers for the new audience. |
| audience | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the safety profile (non-readOnly, non-destructive, non-idempotent), and the description adds a genuinely important trait beyond them: tester generation happens synchronously as a side effect of audience creation, and both the audience and the new testers are returned. It stops short of noting latency, cost, or approval requirements for generated testers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, front-loaded on the create action and its synchronous side effect. Slightly more room could have been spent on the many undocumented parameters without hurting readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be spelled out, but for a 19-parameter mutation tool with nested objects and 11% schema coverage the description is far too thin — an agent has no way to know what most targeting fields mean or how they interact. The synchronous-generation note is the one substantive addition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
19 parameters at 11% schema description coverage means the schema itself documents almost nothing, so the description must carry the load — and it does not. It gestures at filter categories (age, gender, geography, demographics) but leaves roles, languages, industries, techLiteracy, income, maritalStatus, educationLevel, employmentStatus, isReusable, customDetails and more completely unexplained. Far more compensation is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new audience in the caller's Uxia workspace') and immediately discloses the non-obvious secondary effect of synchronously generating AI testers. It also defines what an audience is (targeting filters for recruiting testers), which distinguishes it clearly from read-only siblings like list_audiences.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is strongly implied by 'Create a new audience' and the recruitment framing, but the description never states when to reach for this tool versus alternatives such as list_audiences or set_test_audience, nor does it mention prerequisites like workspace permissions. Adequate but no explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_block_testAInspect
Create an empty draft from a title; then add_task and set_test_audience. Returns the draft revision and app URL. No launch.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark it as a non-read-only, non-destructive, non-idempotent write. The description adds genuinely new behavioral context beyond them: the draft is created empty, it returns the draft revision and app URL, and it does not launch the test. That last point is a useful boundary on a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the create action comes first and the downstream steps follow. The semicolon-joined fragments are terse to the point of slightly clipped grammar, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be detailed, and the description still notes revision and app URL. For a single-parameter setup tool the coverage is close to sufficient, with only the draft-vs-test terminology and sibling routing left thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter with 0% schema description coverage; the length constraints exist but carry no semantic meaning. The description's 'from a title' does tie the title parameter to the draft's naming, adding a little meaning, but nothing about format or rejection behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: create an empty draft from a title. The tool name (create_block_test) and the description's 'draft' framing create mild ambiguity about whether a test is being created, and no sibling (e.g. create_audience, prepare_test_block) is explicitly contrasted, but the core action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It prescribes the follow-up sequence (add_task, set_test_audience) and an exclusion ('No launch'), which implies this is a setup step rather than a launch step. However, it never says when to choose this over create_audience or prepare_test_block, so the routing among siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_draft_advancedADestructiveIdempotentInspect
Execute an advanced draft edit discovered with find_operations. Supply arguments matching its fetched schema, including expectedRevision. Cannot launch or spend credits.
| Name | Required | Description | Default |
|---|---|---|---|
| arguments | Yes | ||
| operation | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=true. The description adds specific, valuable behavioral constraints beyond annotations: it cannot launch or spend credits, and requires an expectedRevision for concurrency control. However, it doesn't detail what an advanced edit can destroy despite the destructive hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero waste. The key constraint (expectedRevision, no launch/spend) is front-loaded after the purpose, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 2 required params with 0% schema coverage, nested objects, and an output schema present, the description provides essential context: discovery workflow, expectedRevision, and execution limits. It doesn't explain the output schema's role, but that is acceptable since the output schema exists to handle return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the two required parameters (operation, arguments). The description compensates by explaining that 'operation' comes from find_operations and 'arguments' must match its fetched schema, and adds a critical expectedRevision requirement likely nested in arguments. However, it doesn't fully specify the operation ID format or argument structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Execute an advanced draft edit') and resource, and explicitly ties it to the discovery tool find_operations, which is a sibling. This clearly distinguishes it from siblings like launch_test and read_advanced.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the prerequisite workflow: the operation must first be discovered with find_operations, and the arguments must match its fetched schema. It also provides clearer exclusions ('Cannot launch or spend credits'), though those are framed behaviorally rather than as alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_operationsARead-onlyIdempotentInspect
Discover advanced Uxia operations. Search query returns up to five summaries, without schemas. Then supply an exact operation name to fetch its schema before using read_advanced or edit_draft_advanced. Discovery never executes an operation.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Find operations by topic: surveys, prototypes, benchmarks, transcripts, navigation, accessibility or construction. | |
| operation | No | Exact name from search; returns its full input schema. |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | |
| next | No | |
| hasMore | No | |
| operations | No | |
| annotations | No | |
| description | No | |
| inputSchema | No | |
| executionTool | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered; the description still adds real behavior beyond them — the capped five-summary search result, the fact that summaries carry no schemas, and the guarantee 'Discovery never executes an operation.' It stops short of describing output shape or behavior on an unmatched name, but with an output schema present that is not required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, then the lookup workflow, then the no-execution guarantee. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter discovery tool with full schema coverage, annotations, and an output schema, the description supplies everything an agent needs: purpose, the two-phase usage pattern, and the safety guarantee. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the per-parameter ground is covered (topic list for query, 'exact name from search' for operation). The description adds sequencing semantics the schema cannot express: query is the entry point, operation is the second-phase lookup keyed to a search result. Useful, though it does not add syntax or format detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Discover advanced Uxia operations') and immediately defines scope: it is a discovery tool, not an executor. An agent can distinguish it from read_advanced or edit_draft_advanced without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit two-step workflow (search to get up to five summaries, then supply an exact operation name to fetch the schema) and names the downstream tools (read_advanced, edit_draft_advanced) that require that schema. The condition selecting each parameter mode is spelled out rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_insightARead-onlyIdempotentInspect
Read a finding by exact ref from list_insights, including evidence. Includes up to three available screenshot previews of referenced frames. Use get_frame_details for additional frames. Images may be unavailable. Use IDs from list_test_blocks to target a block/run; omission uses the documented defaults. Legacy tests also supported.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| runId | No | ||
| testId | Yes | ||
| blockId | No | Select one block from list_test_blocks. Omit to infer a sole relevant block; list_insights defaults to all AI blocks. Omit for legacy tests. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ref | No | |
| data | No | |
| retry | No | |
| runId | No | |
| scope | No | |
| title | No | |
| frames | No | |
| status | No | |
| testId | No | |
| blockId | No | |
| message | No | |
| paywall | No | |
| category | No | |
| nextStep | No | |
| priority | No | |
| relevance | No | |
| userImpact | No | |
| description | No | |
| severityCard | No | |
| totalTesters | No | |
| fixSuggestion | No | |
| businessImpact | No | |
| testerEvidence | No | |
| primaryCategory | No | |
| secondaryCategory | No | |
| technicalComplexity | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuinely useful context: up to three screenshot previews are included, images may be unavailable, and legacy tests are supported. It does not mention pagination or rate limits, but adds real behavioral detail beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with purpose before caveats and alternatives. The 'Legacy tests also supported' clause is slightly orphaned but every sentence carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers previews, image unavailability, and block/default behavior. Gaps remain on runId/testId semantics given the thin schema coverage, but for a read tool it is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, so the description must compensate. It adds meaning for ref (comes from list_insights) and blockId (target a block/run, omit to use defaults, legacy), but says nothing about runId or testId semantics. Partial compensation, so a baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (a finding by exact ref) plus scope (including evidence). It explicitly ties the ref to list_insights and distinguishes itself from get_frame_details, so an agent can pick it apart from siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names an explicit alternative for the adjacent need ('Use get_frame_details for additional frames') and clarifies how blockId targeting works versus omission ('uses the documented defaults'). No explicit when-not-to-use beyond the frame case, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_testBRead-onlyIdempotentInspect
Read test summary and status by ID. Use IDs from list_test_blocks to target a block/run; omission uses the documented defaults. Legacy tests also supported.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| goal | Yes | |
| type | Yes | |
| title | Yes | |
| source | Yes | |
| status | Yes | |
| testId | Yes | |
| version | No | |
| createdAt | Yes | |
| hasResults | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered without the description. The description adds only the legacy-support note and ID sourcing, leaving return/pagination behavior unexplained (though an output schema exists).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and free of padding. The only weakness is that the second sentence's default-omission clause is misleading rather than helpful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering the read-only profile, the description needn't explain returns. Still, for a tool with an ambiguous name amid many get_test_* siblings, it should state how it differs and reconcile the required testId with the 'omission' claim; those gaps leave the agent under-informed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% on a single required parameter, so the description must carry the load. It does identify the ID source (list_test_blocks) and implies a block/run identifier, but the statement that omission 'uses the documented defaults' directly conflicts with testId being required, muddying rather than clarifying the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Read test summary and status by ID'), which is clear on its own. However, it does not differentiate itself from closely named siblings like get_test_progress or get_test_definition, which an agent would need to disambiguate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives partial guidance by pointing to list_test_blocks as the ID source and noting legacy support, which is useful context. But it offers no when-to-use versus the get_test_* siblings, and the 'omission uses the documented defaults' claim is unusable because testId is required by the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_test_definitionARead-onlyIdempotentInspect
Read the current draft, revision, validation and editability before editing. Includes ordered blocks and audience; excludes credential secrets.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so the safety profile is covered; the description adds genuinely new scope information by stating it includes ordered blocks and audience while excluding credential secrets, which tells the agent what data will and will not appear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, with the action and its scope front-loaded and no filler. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the description covers content inclusions/exclusions and the pre-edit timing. The only gap is parameter meaning, which is left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single testId parameter has no documented meaning beyond the 24-hex pattern in the schema. The description offers no explanation of what identifier testId is (test vs block vs launch), so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Read) and enumerates the returned facets (draft, revision, validation, editability), which lets an agent distinguish it from a lighter read like get_test or get_test_progress. It never explicitly names the resource as a test definition, so a small inference is still required.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before editing' implies the intended workflow (read-before-write), which is useful, but no alternative sibling is named and no condition for choosing it over get_test, read_advanced, or edit_draft_advanced is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_test_progressARead-onlyIdempotentInspect
Read run status, progress and credit usage. Omit runId for the latest run.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | No | ||
| testId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | No | |
| url | Yes | |
| runId | Yes | |
| blocks | Yes | |
| status | Yes | |
| testId | Yes | |
| outcome | No | |
| version | Yes | |
| creditsCharged | Yes | |
| currentBlockId | No | |
| creditsRemaining | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare read-only, idempotent, non-destructive, and closed-world behavior, so the safety profile is fully covered. The description adds the scope of returned fields (status, progress, credit usage), which is mild added value; with an output schema present, return-value detail is not required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the resource scope followed by the one parameter rule that matters. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema covering return values and annotations covering safety, the description supplies the essential remaining context: what is read and the runId default. It is only slightly short on testId semantics and failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema alone only shows a 24-hex pattern for testId and a runId ref. The description compensates for runId by explaining that omitting it selects the latest run, but it says nothing about testId's meaning or what happens with an unknown ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource set (run status, progress, credit usage), which is distinct from siblings like get_test or get_test_definition. It is clear what the tool returns, though it never names those siblings to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'Omit runId for the latest run' gives real operational guidance for one parameter, implicitly telling the agent when to supply it and when not to. However, there is no guidance on when to choose this over get_test, show_evidence, or get_insight for run information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
launch_testADestructiveInspect
Launch the previewed revision and spend credits. Requires explicit approval of preview_test_launch costs; maximumCharge is that approved amount. Cost changes require new approval. Locks the definition.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes | ||
| maximumCharge | Yes | ||
| definitionRevision | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | No | |
| url | Yes | |
| runId | Yes | |
| blocks | Yes | |
| status | Yes | |
| testId | Yes | |
| outcome | No | |
| version | Yes | |
| creditsCharged | Yes | |
| currentBlockId | No | |
| creditsRemaining | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description goes beyond them by disclosing the credit spend, the approval-gating rule, and the irreversible side effect of locking the definition. It does not describe failure modes or partial-spend behavior, which keeps it at a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short, front-loaded sentences with no filler. The primary effect (launch and spend credits) leads, followed by the approval precondition and the locking side effect, each sentence carrying a distinct constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. The description covers the prerequisite sibling call, the cost constraint, and the state mutation, which is enough for correct invocation. It leaves failure/rollback behavior and the meaning of the revision parameter unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It gives real semantics for maximumCharge (it is the exact approved amount, with a re-approval trigger if it changes), but says nothing about definitionRevision or testId, leaving two of three required parameters to inference from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (launch) and resource (the previewed revision) plus the material consequence (spends credits, locks the definition). It is clearly distinguishable from the sibling preview_test_launch, which appears as the prerequisite rather than a duplicate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Establishes the workflow prerequisite: costs must first be approved via preview_test_launch, and maximumCharge must equal that approved amount. It also notes that cost changes invalidate the approval and require a new one. It does not state outright when not to use this tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_audiencesList AudiencesARead-onlyIdempotentInspect
List up to 10 audiences in the caller's Uxia workspace. Use this to discover existing audiences before creating a new one, to reference an audience by id in another action, or to summarize the workspace's targeting setup.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| audiences | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds a behavioral fact not in the annotations – the 10-item result cap – which is useful, though it does not say how to page past 10 or what ordering to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The scope and result limit are front-loaded, and the usage scenarios follow compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and a zero-param read tool has little surface to cover. The only gap is that the 10-item truncation is mentioned without guidance on retrieving more, which an agent might reasonably want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool applies. The description correctly does not invent parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List ... audiences"), scopes it to the caller's Uxia workspace, and quantifies the limit (up to 10). An agent can immediately distinguish this retrieval tool from the mutating sibling create_audience.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three concrete contexts for use: discovering audiences before creating one, resolving an audience id for another action, and summarizing targeting setup. It indirectly routes the agent away from create_audience, but never states an explicit exclusion or a when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_insightsARead-onlyIdempotentInspect
List findings ranked by severity. Omit blockId for top insights across all AI task blocks; limit applies globally. Supply blockId for one task. Keep returned block/run IDs with evidence. Use IDs from list_test_blocks to target a block/run; omission uses the documented defaults. Legacy tests also supported.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| runId | No | ||
| sortBy | No | severity | |
| testId | Yes | ||
| blockId | No | Select one block from list_test_blocks. Omit to infer a sole relevant block; list_insights defaults to all AI blocks. Omit for legacy tests. | |
| category | No | ||
| severity | No | ||
| testerName | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| retry | No | |
| runId | No | |
| scope | No | |
| total | No | |
| status | No | |
| testId | No | |
| blockId | No | |
| message | No | |
| paywall | No | |
| insights | No | |
| nextStep | No | |
| totalTesters | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive. The description adds value beyond them: the limit is global when blockId is omitted, returned block/run IDs should be retained for evidence, and legacy tests are supported. These are non-obvious behavioral notes, though it stops short of describing result shape (covered by output schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, but the omission/ID guidance is restated twice ('Omit blockId...' and again 'omission uses the documented defaults'), which adds redundancy without new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description covers the main selection logic, global-limit behavior, and legacy-test support. Remaining gaps are the undocumented filter parameters, which is a minor shortfall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 13% across 8 params, so the description must compensate. It clarifies blockId omission semantics and the global scope of limit, but leaves runId, sortBy, category, severity, and testerName unexplained, so compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List findings ranked by severity') and distinguishes scope from the singular get_insight sibling. It does not name the sibling explicitly, but the list/rank framing is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional usage: omit blockId for cross-block top insights, supply it for a single task, and points to list_test_blocks for valid IDs. No explicit exclusions or named alternative tool, but the when-to-use context is solid.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_test_blocksBRead-onlyIdempotentInspect
List ordered blocks, statuses and exact block/run IDs for result queries.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | No | ||
| testId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | No | |
| url | Yes | |
| runId | Yes | |
| blocks | Yes | |
| status | Yes | |
| testId | Yes | |
| outcome | No | |
| version | Yes | |
| creditsCharged | Yes | |
| currentBlockId | No | |
| creditsRemaining | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds only that results are ordered and include statuses/IDs; it says nothing about pagination, ordering guarantees, or what happens when runId is omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no redundant restatement of the name, which is right for a simple list tool. The closing phrase 'for result queries' is slightly vague but costs nothing structurally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and annotations cover safety. The remaining gap is parameter guidance for a two-parameter tool at 0% schema coverage, which leaves an agent guessing about how to scope the query correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither testId nor runId is explained in the description. The mention of 'block/run IDs' loosely hints at the optional runId, but the required 24-hex testId and the fact that runId is a $ref of testId are never clarified, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List ordered blocks') plus the payload an agent gets back (statuses and exact block/run IDs). It is clear what the tool does, but it never contrasts itself with nearby list siblings such as list_tests or get_test_definition, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The trailing phrase 'for result queries' implies the context in which this tool is useful, but no alternative is named and no precondition (e.g., which test/run must exist first) is stated. Usage is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_testsCRead-onlyIdempotentInspect
List user tests with IDs and status. Use IDs from list_test_blocks to target a block/run; omission uses the documented defaults. Legacy tests also supported.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | ||
| status | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| tests | Yes | |
| total | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and closed-world behavior, so safety is covered. The description adds only the vague 'documented defaults' clause and a block/run targeting statement that does not correspond to any parameter, which is more confusing than informative.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core purpose front-loaded, which is good. However, the 'omission uses the documented defaults' sentence spends space without conveying usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover the safety profile. What is missing is any parameter guidance for three undocumented fields and a coherent explanation of the block/run targeting behavior, which is the main practical gap for an agent trying to call this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters (limit, query, status), so the description carries the full burden and provides no semantics for any of them. The mention of targeting a block/run does not map to any actual parameter, leaving limit's cap, query's scope, and status's enum values unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('List user tests') plus what is returned ('IDs and status'), which is enough to separate it from get_test and get_test_progress. It is slightly muddied by the claim that it can 'target a block/run', which no listed parameter supports.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It references list_test_blocks as a source of IDs but never states when to choose this tool over get_test, get_test_progress, or list_test_blocks itself. 'Omission uses the documented defaults' points at documentation that is not provided, so no real usage condition is established.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_test_blockPrepare missing test detailsARead-onlyIdempotentInspect
Collect up to two missing wording fields for an existing task. Uses forms when supported; otherwise ask returned questions in chat and report that no form appeared. Read-only: save accepted wording with update_task. Never include secrets or treat answers as launch approval.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| testId | Yes | ||
| blockId | Yes | ||
| scenario | No | ||
| stopCondition | No | ||
| expectedRevision | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| saved | No | |
| status | Yes | |
| testId | No | |
| blockId | No | |
| message | No | |
| wording | No | |
| nextTool | No | |
| questions | No | |
| expectedRevision | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only and idempotent, and the description adds substantial beyond that: the form-vs-chat fallback behavior, the instruction to report when no form appeared, the secret-handling guardrail, and the explicit note that answers are not launch approval. This is rich behavioral context an agent could not infer from the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the core action, followed by behavior and guardrails. No filler, though the closing warning sentence is somewhat tangential to the collection mechanics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the behavioral flow is well covered. However, with 0% schema description coverage on 6 parameters, the definition leaves parameter meaning largely unexplained, which is the main completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters (task, scenario, stopCondition, testId, blockId, expectedRevision), and the description only vaguely alludes to 'up to two missing wording fields' without mapping them to the scenario/stopCondition params or explaining required IDs and revision. It does not compensate for the documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Collect up to two missing wording fields for an existing task.' This is clear and actionable, though the tool name 'prepare_test_block' suggests block/test-block preparation while the description frames it around task wording fields, which creates mild ambiguity about the object being operated on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real usage context – use forms when supported, otherwise ask the returned questions in chat and report that no form appeared – and routes to the correct follow-up ('save accepted wording with update_task'). It does not state when not to use this versus siblings, but the operational flow is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_test_launchARead-onlyIdempotentInspect
Preview the saved revision’s sequence, participants, costs and blockers. Present costs for explicit approval before launch_test. Any edit requires a new preview and approval.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes | ||
| definitionRevision | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| testId | Yes | |
| details | No | |
| blockers | Yes | |
| audiences | No | |
| canLaunch | Yes | |
| blockCosts | No | |
| expectedCredits | No | |
| creditsRemaining | No | |
| freeLaunchApplied | No | |
| definitionRevision | Yes | |
| expectedTesterCount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful workflow behavior beyond annotations: edits invalidate the preview and require a fresh approval cycle, and costs must be surfaced for explicit approval. It stops short of describing output shape, but that is covered by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the purpose, then the approval requirement, then the invalidation rule. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and annotations carry the safety profile. The workflow gate (preview → approval → launch, and re-preview after edits) is fully conveyed, leaving only the undocumented parameter meanings as a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required params. The phrase 'saved revision' loosely maps to definitionRevision, but the description never clarifies testId, the 24-hex ID format, or what selecting a specific revision version implies. It does little to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Preview) and enumerates the resource contents it surfaces: sequence, participants, costs and blockers. It is clearly distinguishable from the sibling launch_test, which it explicitly names as the subsequent action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: present costs for approval before launch_test, and re-preview after any edit. This tells the agent both when to call it and when a prior call is invalidated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_advancedARead-onlyIdempotentInspect
Execute an advanced read discovered with find_operations. Supply arguments matching its fetched schema. Cannot edit or launch.
| Name | Required | Description | Default |
|---|---|---|---|
| arguments | Yes | ||
| operation | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | Operation-specific results for a native block test; legacy reads return fields at the top level. |
| retry | No | |
| runId | No | |
| status | No | |
| testId | No | |
| blockId | No | |
| message | No | |
| nextStep | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered elsewhere. The description only restates the negative boundary and the argument-matching requirement; it says nothing about error handling for an unknown operation name or how the fetched schema is validated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the prerequisite, then the exclusion. Every clause carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the dynamic-dispatch nature of the tool is addressed by pointing at find_operations and the fetched schema. The definition is nearly complete for a generic dispatcher, though it omits failure behavior for invalid operations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the 'arguments' parameter is a free-form object, so the schema alone cannot tell an agent how to populate it. The description supplies the missing semantics: 'operation' is a name obtained from find_operations and 'arguments' must conform to the schema fetched for that operation. This is meaningful compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (execute an advanced read), its provenance (an operation discovered with find_operations), and its boundary ('Cannot edit or launch'), which separates it from edit_draft_advanced and launch_test. It stops short of explaining what family of resources these reads cover, so it is clear rather than exhaustive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit prerequisite chain - first find_operations to discover the operation, then pass arguments matching the fetched schema - plus an explicit when-not ('cannot edit or launch'). No named alternative for the read case itself, but the workflow context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_testReview test draftARead-onlyIdempotentInspect
Open one optional draft review card. Use get_test_definition for later reads. Returns draft data without UI support; tool success does not confirm rendering. Editing is revision-checked and never approves launch.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive), so the bar is lower, and the description adds genuinely non-obvious caveats: success does not confirm rendering, returns draft data without UI support, and editing is revision-checked and never approves launch. These are exactly the surprises an agent needs to know, though the interaction between 'editing' and readOnlyHint could be spelled out more.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and the key caveats, with no filler. The telegraphic style ('without UI support', 'revision-checked') is dense but each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description still covers the important behavioral caveats for what is essentially a read of a draft. The main gap is that it never explains the testId input or when to prefer this over preview/launch siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter with 0% schema description coverage, and the description never mentions testId, its ID format, or what draft it addresses. The pattern in the schema is discoverable, but the description adds no meaning beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (open a draft review card) and names the sibling used for subsequent reads (get_test_definition), which helps separate it from get_test, get_test_definition, and get_test_progress. The phrase 'optional draft review card' is somewhat jargon-y about what a 'card' is, but the verb+resource intent is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one routing rule — 'Use get_test_definition for later reads' — but does not say when to open a review card at all, nor how it relates to preview_test_launch, edit_draft_advanced, or launch_test. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_surveyADestructiveIdempotentInspect
Add or replace a Survey block’s questions. Use a new blockId to append, or an existing Survey ID to edit. Other blocks are preserved. Requires current revision. No launch.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes | ||
| blockId | Yes | ||
| questions | Yes | ||
| expectedRevision | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive=true and idempotent=true; the description adds value by scoping the destruction ('Other blocks are preserved'), imposing an optimistic-concurrency prerequisite ('Requires current revision'), and clarifying that this does not launch the test.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short, front-loaded sentences with no filler; the core add/replace semantics come first, followed by mode selection, preservation guarantee, and constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. For a mutation with revision requirements, the description covers safety scope, prerequisites, and side effects, though it omits guidance on how the questions payload is validated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning. It usefully explains blockId (new vs existing) and expectedRevision (must be current), but says nothing about testId or the structure/limits of the questions array (e.g., max 10 items, per-type fields), leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add or replace a Survey block's questions') and immediately disambiguates the two modes via blockId semantics. An agent can distinguish append-vs-edit behavior from the description alone, without needing the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit operating conditions: use a new blockId to append, an existing Survey ID to edit, and it requires the current revision and does not launch. It stops short of naming a sibling alternative, but the when/when-not is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_test_audienceADestructiveIdempotentInspect
Replace shared audience selections and participant counts; null clears. Use existing audience/tester IDs. Preserves blocks. No launch.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes | ||
| audience | Yes | ||
| construction | No | Optional decisions; proposed wording stays separate. Acceptance needs current proposalRevision and matching saved wording. Never launch approval. | |
| expectedRevision | Yes | Current definition.revision. Reread on conflict; never retry blindly. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds real behavioral detail: what is replaced, that null clears the audience, that blocks are preserved, and that no launch occurs. It stops short of describing auth/permission requirements or the response.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three telegraphic clauses with no filler, and the core replace/clear semantics are front-loaded. Dense and efficient, though the clipped style trades some readability for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, revision-guarded mutation with an output schema and annotations present, the description covers the essential behaviors (replace, clear-on-null, preserve blocks, no launch). Concurrency handling is left to the schema's expectedRevision description, which is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, with the schema documenting expectedRevision (revision/conflict) and construction inline. The description adds meaning only for the audience parameter ('null clears', 'use existing audience/tester IDs') and says nothing about construction or expectedRevision, so compensation for the coverage gap is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Replace shared audience selections and participant counts', which clearly distinguishes it from create_audience, list_audiences, add_task, and set_survey. It does not explicitly name a sibling, but the resource and operation are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use existing audience/tester IDs' and 'No launch' imply constraints and hint at the boundary with launch_test, but the description never states when to choose this over siblings like create_audience or set_survey. Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_evidenceShow screenshotsARead-onlyIdempotentInspect
Display an inline screenshot gallery. After reading evidence, reuse the same operation and exact arguments, including test/run/block IDs. Tool success does not confirm that the client rendered the gallery.
| Name | Required | Description | Default |
|---|---|---|---|
| arguments | Yes | Same arguments as the evidence read; never pass image bytes or URLs. | |
| operation | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| imageCount | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive behavior, but the description adds a genuinely useful caveat beyond them: tool success does not confirm the client rendered the gallery. It also clarifies the reuse-the-same-arguments execution model, which is not visible in structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the primary action, followed by the two constraints that matter. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. Combined with the render caveat and argument-reuse instruction, the description gives everything needed to call this display tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema description coverage, the description compensates by narrowing what belongs in arguments (test/run/block IDs, reuse of the read call's exact arguments) and the schema reinforces the 'no image bytes or URLs' constraint. The enum values for operation are left to the schema, so this is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: display an inline screenshot gallery. An agent can distinguish this from the read siblings (get_insight, get_frame_details, get_screen_details) since it renders rather than fetches, though the description never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering context: use it after reading evidence, reusing the same operation and exact arguments. It does not enumerate alternatives or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_taskADestructiveIdempotentInspect
Edit task/scenario/stopCondition wording by blockId. Only supplied fields change; all other settings are preserved. Empty strings clear wording. Requires current revision. Use find_operations for experience or credential changes. No launch.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| testId | Yes | ||
| blockId | Yes | ||
| scenario | No | ||
| stopCondition | No | ||
| expectedRevision | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Open the draft in Uxia. |
| testId | No | |
| editable | No | |
| definition | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: describes partial-merge semantics ('Only supplied fields change; all other settings are preserved'), that empty strings clear wording, and that a current revision is required (optimistic concurrency). These are non-obvious behaviors an agent must know before mutating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five short clauses, front-loaded with the action and scope, then merge semantics, clearing semantics, and prerequisites. No filler; every sentence carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and mutation behavior is well covered. Only testId's role and format go unexplained, leaving a minor hole for a 6-parameter mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry the load, and it does for blockId, task/scenario/stopCondition, and expectedRevision ('Requires current revision'). testId is left unexplained, which is the only gap among six parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edit task/scenario/stopCondition wording by blockId') and scopes it to wording fields. It differentiates from siblings explicitly by naming find_operations for experience/credential changes and implying launch_test is out of scope with 'No launch'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: use this for wording edits, use find_operations for experience/credential changes, and it does not launch. No explicit when-not clause covering other edit paths (e.g. edit_draft_advanced), but the routing guidance is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
22 tool updates
- First observed
add_task - First observed
create_audience - First observed
create_block_test - First observed
edit_draft_advanced - First observed
find_operations - First observed
get_insight - First observed
get_test - First observed
get_test_definition - First observed
get_test_progress - First observed
launch_test - First observed
list_audiences - First observed
list_insights - First observed
list_test_blocks - First observed
list_tests - First observed
prepare_test_block - First observed
preview_test_launch - First observed
read_advanced - First observed
review_test - First observed
set_survey - First observed
set_test_audience - First observed
show_evidence - First observed
update_task
Publisher details
- Operator
- Uxia · Publisher source
- Operator website
- https://uxia.app · Publisher source
- Vendor relationship
- Not applicable
- Documentation
- https://platform.uxia.app/docs · Publisher source
- Trust center
- Not applicable
- Restrictions
- Paid plan required
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.167 npm1MIT
- AlicenseCqualityAmaintenanceCompetitor Monitor AI - MCP server providing AI-powered tools and automation by MEOK AI Labs119 npm37 PyPIMIT
- AlicenseNot gradedqualityBmaintenanceEnables tracking competitor websites, changelogs, blog feeds, and pricing pages with meaningful diffs, classification, and Markdown digests via MCP tools for listing, adding, removing competitors, running checks, and retrieving digests or changes.MIT

industrylens-mcpofficial
AlicenseNot gradedqualityBmaintenanceBrowse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.