Loadster
Server Details
Load testing and synthetic monitoring platform: test with Playwright, Browser Bot, or Protocol Bots.
- Status
- Healthy
- Uptime
- 96.8% over 21 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- loadster/loadster-mcp
- GitHub Stars
- 0
TDQS
Scored across 55 tools
Every tool targets a distinct resource and action, and pairs like list_scenarios/get_scenario or get_play_status/get_step_detail are clearly separated by whether they return summaries or full detail. Despite 55 tools, there are no overlapping or mutually confusable tool purposes.
All tools use consistent lowercase snake_case verb_noun naming: create_*, get_*, list_*, update_*, delete_*, play_*, validate_*, etc. Even less common verbs like append_, disable_, duplicate_, and restore_ follow the same predictable pattern.
At 55 tools, this is an extreme tool count, falling well beyond the already 'too many' threshold. Each tool may be legitimately scoped, but the sheer size of the surface would overwhelm agents and cause decision fatigue, making the set inappropriate for typical agent workflows.
The tool set provides comprehensive CRUD/lifecycle coverage for scripts, datasets, scenarios, monitors, and includes deep reporting and debugging support. The main gaps are deliberate: no project creation, no monitor enabling, and no load-test launching/stopping, which are left to the Dashboard and could be dead ends for a brand-new team.
Available Tools
55 toolsappend_dataset_rowsAppend Dataset RowsAInspect
Add rows to the end of an existing dataset, sending only the new rows — the rows already in the table stay as they are, so there is no need to read them first. Every row must have as many columns as the table already has; a row that doesn't is rejected and nothing at all is written. Appending to an empty dataset takes the column count from the first row.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | The rows to add to the end of the table (data rows only — no header row). Each row must have the same number of columns as the dataset. | |
| datasetId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable behavioral details beyond the annotations: if any row has the wrong column count, the whole write is rejected and nothing is written, and appending to an empty dataset infers column count from the first row. This atomic failure behavior is not expressed in the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is composed of three purposeful sentences: the first states the action, the second explains the key behavioral guarantee, and the third handles an edge case. Every sentence adds necessary information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter mutation tool with no output schema, the description covers the core call mechanics, validation rules, failure behavior, and an important edge case. An agent has enough context to invoke the tool correctly without ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, but the description compensates for the most important parameter, rows, by specifying that rows are data-only, must match the existing column count, and are rejected otherwise. projectId and datasetId are not described in detail, but their names plus the 'existing dataset' context make their meaning obvious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource combination—'Add rows to the end of an existing dataset'—which clearly distinguishes it from create, update, and delete operations. The phrase 'existing dataset' also prevents confusion with dataset creation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when appending is appropriate: only new rows are sent, existing rows are untouched, and there is no need to read the table first. It does not explicitly name an alternative tool like update_dataset, but the usage context is clear enough for an agent to choose it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_datasetCreate DatasetAInspect
Create a new dataset in a project. Loadster treats every row as data and addresses columns by zero-based index — a header row would be fed to bots like any other row, so don't add one when authoring values. Unless the user explicitly asked for a dataset, confirm with them before creating one.
| Name | Required | Description | Default |
|---|---|---|---|
| dataset | Yes | The dataset name and table of values (data rows only — no header row). | |
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only state readOnly=false, idempotent=false, etc.), the description warns about a critical gotcha: header rows are treated as data and will be fed to bots, so they should not be added. It also discloses the confirmation requirement, adding genuine behavioral context that the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first states the core action, the second explains a non-obvious data formatting constraint, and the third provides a user-confirmation rule. The most important operational warning is front-loaded and clearly stated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the nested schema and no output schema, the description covers the main non-obvious aspects: header rows, zero-based indexing, and confirmation before creation. It is slightly incomplete in not mentioning what happens on success or what the return value is, but this does not block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful detail for the 'values' parameter by explaining zero-based indexing and the no-header-row rule, which goes beyond the schema's minimal 'data rows only' note. However, 'projectId' is not described in the schema or the description, so parameter semantics are only partially enriched.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Create a new dataset in a project.' It also clarifies the essential domain behavior that every row is treated as data with zero-based column indexing, which helps distinguish this from generic create operations. This is precise and immediately actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear precondition: confirm with the user before creating a dataset unless explicitly asked. However, it does not mention when to prefer alternatives like append_dataset_rows or update_dataset, so routing between sibling tools is only implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_monitorCreate MonitorAInspect
Create a Monitor in a project — always disabled, and nothing on this surface can enable one. An enabled Monitor consumes fuel on its schedule, unattended, so enabling is the user's act in the Dashboard, at the link the returned summary carries. Read the 'monitoring' topic before choosing a frequency or thresholds.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The monitor name. | |
| tags | No | Tag names to organize the Monitor. Prefer the names other Monitors already carry, as list_monitors reports them — a name that doesn't exist yet becomes a new tag, so don't invent tags the user didn't ask for. Omit for no tags. | |
| scriptId | Yes | The id of the Script in this project the Monitor plays each Cycle. | |
| engineIds | Yes | The ids of the monitoring locations the Cycles run from, as list_monitoring_locations reports them; at least one is required. Monitoring locations are their own set, not the load test regions list_engines offers, and an unknown id is refused with an error naming the valid ones. | |
| frequency | Yes | How often a Cycle runs, in milliseconds: 60000 (one minute) to 86400000 (one day). | |
| projectId | Yes | ||
| thresholds | No | The acceptable-metric limits a Cycle is judged against, each optional within the object: responseTimeAverageMin/Max, responseTimeTotalMin/Max, cycleDurationMin/Max, timeToFirstByteMax, firstContentfulPaintMax, largestContentfulPaintMax, and totalBlockingTimeMax in milliseconds; cumulativeLayoutShiftMax as the unitless Web Vitals score; performanceScoreMin as the 0-100 floor. A limit left out is not checked; omit the whole object to check none. | |
| failureThreshold | No | Consecutive failed Cycles before an Incident opens (1 to 128; default 1). | |
| recoveryThreshold | No | Consecutive passing Cycles before an open Incident closes (1 to 128; default 1). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing that API-created monitors are always disabled, that nothing on this surface can enable them, that enabled monitors consume fuel on a schedule, and that the returned summary carries the Dashboard link for enabling. These are non-obvious behavioral traits an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place: the first states the core behavior, the second explains the fuel consequence and where enabling happens, and the third directs the agent to prerequisite knowledge. No filler or repeated schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 9-parameter tool with a nested thresholds object and no output schema, the description covers the key cross-cutting behaviors, mentions the important returned artifact (the link in the summary), and points to the monitoring topic for nuanced choices. Combined with the detailed schema descriptions, nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 89%, so the input schema already documents the parameters thoroughly. The description's only parameter-related contribution is advising the agent to read the 'monitoring' topic before choosing frequency or thresholds, which adds a small layer of guidance but does not materially improve parameter understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — 'Create a Monitor in a project' — and immediately adds the defining qualifier that the monitor is always disabled, which clearly distinguishes it from update_monitor, delete_monitor, and disable_monitor. This leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context that this tool creates monitors in a disabled state and that enabling can only happen through the user's Dashboard act, which prevents an agent from trying to enable a monitor after creation. It does not explicitly name sibling alternatives like update_monitor, but the context is sufficient to guide correct use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_scenarioCreate ScenarioAInspect
Create a scenario in a project — saving costs nothing and starts nothing, so hand the user the dashboard link and let them launch it. The save is refused for a loadEngineId list_engines doesn't offer, a stage shorter than the minimum, or more bots or populations than the account allows, and the refusal names the limit. The summary's launchBlockers are reasons the user could not launch it right now; they never stop the save, so tell the user about a blocker instead of rewriting the scenario to avoid it. The 'load-test-scenarios' topic covers what the stages and ramps actually do.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The scenario name. | |
| projectId | Yes | ||
| populations | Yes | The Populations, in order. Each needs a name, the scriptId of a script in this project, and stages — { target, duration, ramp } segments that run for duration milliseconds and reach target bots by the end, each starting from where the previous stage left off. Optional per population: id (send back the id get_scenario reported to keep a population's identity), loadEngineId (an id from list_engines; defaults to the account's default region), iterations, iterationsPerUser, aggressionMultiplier, and playbackOptions with connectionBps (bits per second), hostnameOverrides, and variableOverrides. Read get_documentation topic 'load-test-scenarios' for what stages and ramps do; list_engines carries the account's stage bounds and bot caps. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reveals important behavioral traits beyond the annotations: saving has no cost and starts nothing, refusals occur for specific invalid inputs, refusal messages name the limit, and launchBlockers do not prevent saving. Since annotations are all false/uninformative, the description carries the burden and does so richly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences carry a large amount of actionable information with no filler. The most important side-effect statement is front-loaded, followed by validation semantics and user-handling guidance. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with no output schema, the description covers side effects, validation failures, the response summary's launchBlockers, and a pointer to relevant documentation. However, it does not explain how to source or validate projectId, nor name the exact response field for the dashboard link, leaving minor gaps an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents populations and stages in detail, but top-level projectId has no description and name only has a trivial one. The description adds meaningful operational constraints: loadEngineId must come from list_engines, stages must meet minimum duration, and bot/population counts are subject to account limits. This compensates for schema gaps, though it does not enumerate every parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Create a scenario in a project,' a specific verb and resource that clearly identifies the tool's function. The resource 'scenario' is distinct from sibling create_* tools, so an agent can select it without opening the schema. The added context about saving not starting anything further sharpens the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: saving is cheap and inert, so the agent should hand the user the dashboard link and let the user launch the scenario. It also instructs how to handle launch blockers. However, it does not explicitly name alternative tools or state when *not* to use this tool, so it stops short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_scriptCreate ScriptAInspect
Create a new script in a project. Read the relevant get_documentation topics and use the schema tools first to build valid commands for the bot type. Like every write on this surface, returns a summary — the server-assigned facts and a dashboard link — never the object you sent; the matching get tool reads it back.
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | The script name, type (HTTP, BROWSER, or PLAYWRIGHT), commands, and variables. | |
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only provide flags and no return information), the description discloses a key behavioral trait: every write returns a summary with server-assigned facts and a dashboard link, not the object sent, and the matching get tool must be used to read it back. This adds useful behavioral context not available from annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences each earn their place: the first states the action, the second provides necessary preparatory guidance, and the third explains the return contract. The most important information is front-loaded and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the nested script object, enum type, two required parameters, and no output schema, the description covers the essential workflow: consult docs/schema, create, expect a summary, and read back via get. It does not detail every nested field requirement, but that belongs to the schema and referenced schema tools, so the description is reasonably complete for a create operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: the script parameter has a description listing name, type, commands, and variables, while projectId has no description. The tool description advises using schema tools to build valid commands, which indirectly supports parameter usage, but it does not add direct explanations for individual parameters beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Create a new script in a project,' which names a specific verb, resource, and scope. It is clear that this tool is for creating rather than updating, deleting, or importing scripts, though it does not explicitly name sibling alternatives as a differentiator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives actionable preconditions: read get_documentation topics and use schema tools first to build valid commands. However, it does not explicitly state when to choose this tool over alternatives like import_script, duplicate_script, or update_script, leaving usage timing mostly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_datasetDelete DatasetBDestructiveInspect
Delete a dataset from a project.
| Name | Required | Description | Default |
|---|---|---|---|
| datasetId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true and readOnlyHint=false, so the description adds no behavioral depth beyond the fact that a dataset is removed. It does not disclose whether deletion is permanent, whether related data is cascaded/deleted, or what happens to dependent artifacts. Since annotations carry the destructiveness signal, a low but non-contradictory score applies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence that front-loads the action and object; there is no redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive two-parameter tool with no output schema, the description is under-specified: it omits expected effects, failure conditions, and any post-delete behavior. The annotation indicates destruction, but the text alone would not tell an agent what will happen when the tool succeeds.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, but it only restates the object-container relationship ('a dataset from a project') that parameter names already imply. It gives no guidance on required formats, constraints, or the meaning of the two IDs beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Delete'), a specific object ('dataset'), and the containing scope ('from a project'), which is enough to distinguish it from sibling delete tools for monitors, scenarios, scripts, and assets. It is not merely a restatement of the name because it adds project scoping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool's intended use is clear from the phrasing, but the description gives no explicit guidance about when to choose it over alternatives or any preconditions/exclusions. Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_monitorDelete MonitorADestructiveInspect
Delete a Monitor from a project. Its Incidents are deleted with it and it is removed from any status pages; the Script it played is left alone. To stop a Monitor running without losing it, use disable_monitor instead.
| Name | Required | Description | Default |
|---|---|---|---|
| monitorId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description specifies exactly what gets destroyed (Incidents, status-page membership) and what survives (the Script). This is valuable behavioral context that annotations alone do not provide, and it does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each with a distinct job: state the action, list side effects, point to the alternative. There is no filler or redundant restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive delete tool with a simple schema, the description covers the key behavioral concerns: permanent deletion, cascading removal, retention of the Script, and the safer alternative. It does not mention success/failure response semantics, and there is no output schema to cover that gap, so a small omission remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It only says 'from a project,' which weakly maps to projectId and monitorId, but it provides no additional detail about formats, relationships, or required preconditions. The description adds minimal meaning beyond the parameter names already visible in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Delete a Monitor from a project.' It also distinguishes itself from the sibling disable_monitor by explicitly framing deletion as permanent removal and naming the alternative. This makes the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence is an explicit when-not-to-use rule: 'To stop a Monitor running without losing it, use disable_monitor instead.' This gives the agent a clear decision boundary between delete_monitor and its closest sibling, with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_scenarioDelete ScenarioADestructiveInspect
Delete a scenario from a project. The scripts its populations referenced are left alone.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| scenarioId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds a valuable behavioral detail beyond that: the scripts referenced by the scenario's populations are left untouched, so destruction is scoped to the scenario only. This helps the agent understand the side-effect boundary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence with no filler. It front-loads the action and adds the important side-effect clarification immediately, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter delete operation with destructive annotations already present, this description is largely sufficient. It names the action, scope, and the key side-effect boundary. It does not cover error behavior or irreversibility, but those are less critical given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only provides minimal implicit semantics: 'from a project' hints that projectId identifies the project and scenarioId identifies the scenario. It does not explain ID formats, required relationships, or any validation constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Delete'), the resource ('a scenario'), and the scope ('from a project'). This distinguishes it from sibling delete tools like delete_dataset, delete_monitor, and delete_script by naming the specific resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this when you need to remove a scenario from a project. It does not explicitly name alternatives or say when not to use it, but the context is unambiguous enough for an agent to select it over the related delete_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_scriptDelete ScriptADestructiveInspect
Delete a script from a project.
| Name | Required | Description | Default |
|---|---|---|---|
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation already declares destructiveHint=true, and the description merely restates the destructive action without adding behavioral context. It does not mention irreversibility, cascading effects on associated assets, or whether deletion requires special permissions. No extra value beyond the annotation is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-formed sentence with no filler. It immediately states the action, object, and scope, making it easy for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple delete operation with two string parameters and a destructiveHint annotation, the description covers the core necessities. It could be enhanced by noting permanence or lack of idempotency, but the annotation and schema already carry much of that weight, and the tool is simple enough that this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It maps semantics to both parameters: 'a script' corresponds to scriptId and 'from a project' corresponds to projectId, indicating the parent-child relationship. This is useful but minimal; it adds no guidance on where to find these IDs or any value constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Delete'), a specific resource ('a script'), and the containing scope ('from a project'). This unambiguously differentiates it from sibling tools like delete_script_asset, delete_dataset, and delete_monitor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the basic use case obvious: when you want to delete a script within a project. However, it offers no explicit guidance on when not to use it, such as distinguishing between deleting a script and deleting a script asset, or noting any prerequisites like ownership or permissions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_script_assetDelete Script AssetBDestructiveInspect
Delete a file attached to a script.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | The asset path. | |
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already carry destructiveHint=true, so the destructive nature is covered. The description adds a useful nuance: it deletes a file attached to a script, not the script itself. However, it does not disclose effects like irreversibility, failure behavior if the path is missing, or permission requirements beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that states the core action without fluff. It is appropriately short for a straightforward delete operation and avoids repeating the title verbatim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-string-parameter delete operation with no output schema, the description is minimally adequate but not complete. It lacks usage guidance, parameter clarification beyond the schema, and any note about consequences or prerequisites. The destructiveHint annotation covers one behavioral aspect, but overall context is thinner than ideal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%; only 'path' has a schema description. The tool description does not compensate by explaining projectId, scriptId, or path in more detail. 'A file attached to a script' vaguely suggests the identifiers' roles, but the description does not clarify what path means or how it relates to the asset, leaving critical parameter semantics under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Delete') and resource ('a file attached to a script'), clearly distinguishing it from sibling delete_script, which deletes the entire script. It also aligns with the script-asset family (get/list/put_script_asset) without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance about when to use this tool versus alternatives. It does not mention that delete_script should be used for removing the whole script, or that put_script_asset could be used to replace an asset. The intended use is inferable from the name, but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
disable_monitorDisable MonitorAIdempotentInspect
Disable a Monitor, stopping its scheduled Cycles and their fuel burn; disabling closes the Monitor's open Incidents. Idempotent, and the only enabled-state transition this surface offers — re-enabling is the user's act in the Dashboard, at the link the returned summary carries.
| Name | Required | Description | Default |
|---|---|---|---|
| monitorId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark idempotence and non-read-only; the description adds the operational effects: scheduled Cycles stop, fuel burn stops, open Incidents close, and the returned summary contains the Dashboard link for re-enabling. This goes well beyond the structured hints and matches the idempotentHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence, front-loaded with the action and then explaining consequences and constraints. Every clause adds distinct information; there is no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the operation's non-obvious effects and tells the agent what the returned summary carries (the re-enable link), partially compensating for the missing output schema. It does not describe the full return shape or error conditions, so it is not fully complete, but for a simple idempotent disable action the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain `projectId` or `monitorId` beyond the word 'Monitor.' The parameter names are fairly self-evident, but with no descriptions in the schema and no parameter explanations in the text, the agent gets no extra semantic aid for constructing a valid request.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase 'Disable a Monitor' pairs a specific verb with a resource, and the description adds precise consequences: 'stopping its scheduled Cycles and their fuel burn' and 'closes the Monitor's open Incidents.' It also distinguishes itself from siblings by declaring it 'the only enabled-state transition this surface offers,' so it is clearly not a generic update or delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-not: 're-enabling is the user's act in the Dashboard,' telling agents not to expect an API re-enable via this or any sibling. 'The only enabled-state transition this surface offers' clarifies that this is the single place to perform disabling, making the use case unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_scriptDuplicate ScriptAInspect
Duplicate a script, optionally with a new name and into another project the token can access. Returns the copy's summary.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Optional name for the copy; a name is generated when omitted. | |
| scriptId | Yes | The id of the script to duplicate. | |
| projectId | Yes | The project id of the source script. | |
| destinationProjectId | No | Optional destination project id; defaults to the source project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already indicate this is a non-read-only, non-idempotent, non-destructive operation. The description adds meaningful behavior beyond those annotations by specifying that the destination must be 'another project the token can access' and that the copy's summary is returned. This gives useful operational context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action and then adds the key optional behaviors. Every clause earns its place, and there is no redundant restatement of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately includes that the operation 'Returns the copy's summary.' The tool is simple enough that the description plus the 100%-covered schema covers the essential behavioral and parameter information. A minor gap is that it does not mention whether related assets or revisions are copied, but this is not critical for a basic duplicate operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds at least one meaningful nuance beyond the schema by noting the destination project must be accessible by the token, which is a permission-relevant detail for destinationProjectId. It also implicitly maps 'new name' to the name parameter, but the schema already documents these fields well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the operation as 'Duplicate a script' with an optional new name and optional destination project, which is distinct from creating or importing a script. It is a specific verb+resource pair with enough detail to differentiate it from sibling tools like create_script and import_script.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when an existing script needs to be copied, optionally renamed, or moved to another project. It does not explicitly name alternatives or state when not to use it, but the duplicate semantics are clear enough for an agent to select it appropriately among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_command_schemaGet Command SchemaARead-onlyIdempotentInspect
Get the recursive schema for one command type, so you can build a valid command object. Some fields carry a 'notes' string describing accepted keys that aren't reflectable, such as element action options and locator attributes.
| Name | Required | Description | Default |
|---|---|---|---|
| commandType | Yes | The command type, e.g. 'http', 'navigate', or 'element'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| kind | Yes | |
| type | Yes | |
| items | No | |
| notes | No | |
| fields | No | |
| variants | No | |
| allowedValues | No | |
| discriminator | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a safe, read-only, idempotent operation. The description adds useful behavioral detail beyond the annotations: the schema is recursive, and some fields include a 'notes' string covering non-reflectable accepted keys. This helps the agent understand what the returned schema will contain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, with the core purpose front-loaded and the second sentence adding genuinely useful detail about non-reflectable keys. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with an output schema, the description is complete. It tells the agent what the tool does, when to use it, and what to expect in the response, including a notable caveat about 'notes' strings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the input schema already describes the single parameter with an example. The description does not add new parameter-level semantics, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get'), names the resource ('recursive schema for one command type'), and states the intended outcome ('build a valid command object'). This clearly distinguishes it from sibling tools like list_command_types and get_variable_schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the use case: when you need to construct a valid command object for a single command type. It does not explicitly name alternatives or state when not to use it, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_datasetGet DatasetARead-onlyIdempotentInspect
Get a single dataset's shape — name, row and column counts, dashboard link — plus a leading window of its rows: the first five in table order by default, more with offset and limit. Compare rowOffset and the rows returned against rowCount to see how much you are not looking at. Loadster serves every row to bots, but a user's imported data may still start with a header-looking row — don't assume either way; ask before treating or removing it.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many rows to return (default 5, maximum 100). Page through a larger table by raising offset. | |
| offset | No | Zero-based index of the first row to return (default 0). | |
| datasetId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a safe read-only, idempotent operation. The description adds valuable behavior beyond annotations: default first five rows in table order, offset/limit pagination, the rowOffset/rowCount comparison, and the caution about header-looking rows. This materially helps the agent handle returned data correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct information: what is returned, how pagination works, and a critical data caveat. The description is front-loaded with the primary purpose and stays compact without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes responsibility for explaining return values and does so well: shape fields, row window, rowOffset/rowCount, and header-row caution. It provides enough detail for an agent to invoke the tool and interpret results correctly, including pagination and edge-case handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions cover limit and offset, and the description reinforces their default behavior. Required projectId and datasetId have no schema descriptions, but their meaning is clear from names and the tool's purpose. The description does not add substantial new semantics for the parameters beyond what the schema already provides, and 50% coverage is only partially compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description names the exact operation: retrieving a single dataset's shape and a leading window of rows. It clearly differentiates from sibling list_datasets by emphasizing 'single dataset' and enumerates the returned fields (name, row/column counts, dashboard link, rows).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: when you need one dataset's shape and initial rows, with pagination guidance. It does not explicitly name alternatives or say when not to use it, but the context is unambiguous enough for an agent to choose it over list_datasets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_documentationGet DocumentationARead-onlyIdempotentInspect
List the Loadster manual's topics with their keys and URLs, or read up to 3 topics' text by key.
| Name | Required | Description | Default |
|---|---|---|---|
| topics | No | Up to 3 topic keys, e.g. 'dynamic-datasets' or the nested 'playwright-scripts/bot'. Omit to list every topic. | |
| maxLength | No | Characters to return per topic. Defaults to 25000. | |
| startIndex | No | Character offset to resume one topic from, taken from a previous call's nextStartIndex. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds the dual list/read behavior and the up-to-3-topics constraint, which is useful, but it does not disclose additional behaviors such as pagination details or what happens with invalid topic keys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the primary action and clearly separates the two modes with 'or'. Every word earns its place, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with no output schema, the description sufficiently covers the main return types: keys and URLs for the list mode, and text for the read mode. It does not mention nextStartIndex or maxLength defaults, but those are fully documented in the parameter schema, so the description plus schema is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description reinforces the 'by key' semantics and the 3-topic limit, but adds minimal meaning beyond what the schema provides. A baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs and a clear resource: 'List the Loadster manual's topics with their keys and URLs, or read up to 3 topics' text by key.' It clearly distinguishes this tool from sibling get_* tools by targeting the Loadster manual specifically, and it defines two distinct modes of operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for use: list topics when no topics are specified, or read up to 3 topics by key. It does not explicitly name alternatives or exclusions, but no other sibling tool covers the Loadster manual, so the usage context is sufficiently clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exampleGet Example ScriptARead-onlyIdempotentInspect
Get a small, valid worked example script for a bot type to use as a starting point.
| Name | Required | Description | Default |
|---|---|---|---|
| botType | Yes | The bot type: HTTP, BROWSER, or PLAYWRIGHT. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds qualitative details about the output ('small, valid worked example') and the bot type scope, which is useful but not extensive. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One succinct sentence communicates the action, resource, qualifiers, and intended use without any filler. Information is front-loaded and every word contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one enum parameter and no nested objects, the description adequately conveys what the tool returns and why. It could be more explicit about the exact response format, but with annotations covering safety and the description covering purpose, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single parameter botType fully, including an enum and a clear description. The tool description adds no additional parameter-level information, but with 100% schema coverage the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Get' and a specific resource 'example script for a bot type', and adds the purpose 'as a starting point'. It is clearly not the same as get_script, though it does not explicitly name the sibling it differs from; the word 'example' signals the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'to use as a starting point' implies the tool is meant for bootstrapping new scripts, which is a usage context. However, it provides no explicit guidance on when to prefer this over alternatives like get_script or get_scripting_api, and no exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_incidentGet IncidentARead-onlyIdempotentInspect
Get a single Incident: when it opened and closed, every notification that went out with its delivery outcome, the status events recorded on it, and the ids of the Cycles that opened and closed it. To reconstruct what happened, read those Cycles with get_monitor_cycle_detail, and cross-read the Monitor's configuration and its Script so the diagnosis names the concrete step, URL, or threshold at fault rather than a vague symptom. The 'monitoring' topic covers advising on thresholds or noise.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| incidentId | Yes | The incident id, as reported by list_incidents. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful behavioral context by listing what the call returns, including notification delivery outcomes and status events, and by framing it as a read that typically requires follow-up reads. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first two sentences are information-dense and front-loaded with valuable details. However, the final sentence about the 'monitoring' topic covering thresholds or noise is off-topic for invoking get_incident and does not earn its place, making the description less economical than it could be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the return categories and prescribing the reconstruction workflow. It is complete enough for an agent to understand what it will get and what to do next. The minor off-topic sentence does not undermine completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents incidentId as 'as reported by list_incidents', and projectId is left to its name. The description does not add parameter-level meaning beyond mentioning that it fetches a single Incident. With 50% schema coverage, this is adequate but not enhanced by the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool gets a single Incident and enumerates the exact data returned: open/close times, notifications with delivery outcomes, status events, and the IDs of Cycles that opened and closed it. The 'single' wording distinguishes it from list_incidents, and the referenced get_monitor_cycle_detail helps differentiate it from related detail tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit follow-up guidance: read the Cycles with get_monitor_cycle_detail and cross-read the Monitor configuration and Script to reach a concrete diagnosis. It does not explicitly say 'use this instead of list_incidents', but the single-incident framing plus the schema hint that incidentId comes from list_incidents makes the intended context clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_load_test_reportGet Load Test ReportARead-onlyIdempotentInspect
Get a finished Load Test's Test Report — aggregate totals, the error and advice picture behind the Dashboard's advice tile, and per-Population outcomes; an unfinished test answers with an error naming its status. The advice buckets and detection flags are heuristics, not a diagnosis of this test: where the totals, percentiles, per-URL counts and Populations contradict a bucket or a flag, or show something it never names, say so and follow the data. Before advising the user, read the get_documentation topic 'analyzing-test-results' and cross-read the Scenario and the played Scripts, so your advice names concrete elements — a specific URL, Population, bot count, or script step — instead of repeating the error buckets back.
| Name | Required | Description | Default |
|---|---|---|---|
| testId | Yes | The load test id, as reported by list_load_tests. | |
| projectId | Yes | ||
| errorUrlLimit | No | How many URLs to return in errorsByUrl, worst first (default 10, no maximum). There is no offset, so read the tail by raising this. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| name | Yes | |
| notes | No | |
| advice | No | |
| maxBots | Yes | |
| startedAt | No | |
| totalHits | Yes | |
| finishedAt | No | |
| scenarioId | No | |
| totalPages | Yes | |
| errorsByUrl | Yes | |
| populations | Yes | |
| totalErrors | Yes | |
| dashboardUrl | Yes | |
| errorUrlNote | No | |
| scenarioName | No | |
| stoppedEarly | Yes | |
| playedScripts | Yes | |
| responseTimes | Yes | |
| totalErrorUrls | Yes | |
| totalIterations | Yes | |
| avgHitsPerSecond | Yes | |
| maxHitsPerSecond | Yes | |
| avgPagesPerSecond | Yes | |
| maxPagesPerSecond | Yes | |
| totalBytesTransferred | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior, so the description goes well beyond that by disclosing that unfinished tests return an error with status, and that advice buckets and detection flags are heuristics rather than diagnoses. It also instructs the agent to trust the underlying data over the buckets and to state contradictions, which is valuable behavioral context beyond any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded: the first sentence gives the core purpose and key output categories. The second and third sentences add important interpretation guidance. The final sentence is somewhat long and mixes post-invocation workflow guidance with tool usage, so it is not maximally concise, but every sentence contributes meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description does not need to enumerate return fields. It covers error behavior for unfinished tests, the heuristic nature of advice buckets, and the prerequisite reading needed to use the result responsibly. Combined with the annotations and schema, nothing critical is missing for an agent to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: testId and errorUrlLimit are documented in the schema, but projectId is not described anywhere. The tool description does not add parameter-level meaning, though it does clarify the testId context by tying the report to finished tests. Since the description does not compensate for the undocumented projectId, this is adequate but not strong.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: 'Get a finished Load Test's Test Report.' It then enumerates the concrete parts of that report (aggregate totals, error/advice picture, per-Population outcomes) and even defines behavior for unfinished tests. This is enough to distinguish it from the many sibling get_* tools without needing to name one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes that this tool is for finished load tests and that unfinished tests will respond with an error naming the status. It also gives explicit follow-up guidance to read get_documentation's 'analyzing-test-results' topic and cross-read the Scenario and Scripts before advising. It does not name a specific sibling alternative, but the resource is unique enough that this is not a significant gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_monitorGet MonitorARead-onlyIdempotentInspect
Get a single Monitor's full configuration and dashboard link. Configuration only — a Monitor's steps are its Script's commands, so read them with get_script. Whether a Monitor runs is the user's call: enabling happens in the Dashboard at the link, never through this surface. The 'monitoring' topic covers what a Monitor's configuration should be.
| Name | Required | Description | Default |
|---|---|---|---|
| monitorId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is established. The description adds meaningful context by limiting the scope to configuration only, clarifying that steps are not included, and explaining that enabling is not a side effect of this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first two sentences are focused and front-loaded, but the final sentence about the 'monitoring' topic is vague and does not clearly earn its place. It adds little actionable guidance and slightly dilutes the otherwise tight structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple getter with two obvious parameters and strong annotations, the description conveys what is returned (configuration and dashboard link), what is not returned (steps), and how enabling is handled elsewhere. Lacking an output schema, it could specify return structure more explicitly, but the core invocation context is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description bears the burden of explaining parameters. It does not mention monitorId or projectId at all, nor does it explain their formats or relationships. The parameter names are somewhat self-explanatory, but the description provides no compensating detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves a single Monitor's full configuration and dashboard link. It explicitly contrasts with get_script for reading steps, distinguishing it from a key sibling tool and preventing mis-selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names get_script as the alternative for reading a Monitor's steps and states that enabling a monitor is done through the Dashboard, never through this surface. This gives clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_monitor_cycle_detailGet Monitor Cycle DetailARead-onlyIdempotentInspect
Get one Cycle in full: status, failure message, the engine it ran from, timestamps, and every metric the Monitor's thresholds judge. Its scriptRunId is a played run like any other — pass it with the project id to get_play_status for per-step results and get_step_detail for the exact request and response that failed — and it names the scriptId it checked, so cross-read the Script and point advice at a concrete step or URL. Works with a cycle id from list_monitor_cycles or from an Incident's cycle references. The 'monitoring' topic covers what the thresholds should be.
| Name | Required | Description | Default |
|---|---|---|---|
| cycleId | Yes | The cycle id, as reported by list_monitor_cycles or an Incident's cycle references. | |
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context beyond the schema: it explains that scriptRunId behaves like a played run, that the cycle names a scriptId for cross-reading, and how the detail connects to incident cycle references. It doesn't disclose things like pagination or response shape, but there is no output schema and the key relational traits are well described; a 4 is warranted given the strong context added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences and front-loads the primary purpose ('Get one Cycle in full') before moving to related-tool routing. Every sentence adds either outcome detail or cross-tool guidance, and it avoids fluff. The only structural weakness is that it's a dense run-on sentence in the middle, but it remains readable and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only two-parameter detail tool with annotations covering safety, the description is nearly complete: it names the fields returned, explains cycleId provenance, and routes to sibling tools for deeper drill-down. It doesn't specify the exact output shape or mention what happens when a cycle is missing or in progress, and there is no output schema, but the essential usage context for an agent is present. A 4 reflects that it's strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents cycleId with the same source note ('as reported by list_monitor_cycles or an Incident's cycle references'), so the description's one parameter detail is redundant. projectId has no schema description and the description doesn't explain what the project id is for or how to obtain it. Schema coverage is 50%, so the description only partially compensates; it clarifies the source of cycleId but leaves projectId unexplained, so a 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('one Cycle in full') and enumerates exactly what fields are returned: status, failure message, engine, timestamps, and metrics. It also distinguishes this tool from siblings like list_monitor_cycles, get_play_status, and get_step_detail, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: with a cycle id from list_monitor_cycles or an Incident's cycle references. It also gives concrete routing guidance to alternatives—pass scriptRunId with project id to get_play_status for per-step results and get_step_detail for the exact failed request/response—and points to the 'monitoring' topic for threshold context. This is explicit when-to-use and when-to-use-other guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_monitoring_summaryGet Monitoring SummaryARead-onlyIdempotentInspect
Get the Monitoring Summary for a project: the aggregate health picture across its Monitors over a period, defaulting to the last 7 days, with per-Monitor and per-day aggregates for trends (times in seconds; performance score 0-100). Drill down instead of asking for more: list_incidents for what opened and closed in the period, list_monitor_cycles for one Monitor's individual Cycles. Read the 'monitoring' topic before interpreting uptime or health for the user.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Start of the period, ISO-8601 (2026-08-19 or 2026-08-19T00:00:00Z). Defaults to seven days before until. | |
| until | No | End of the period, ISO-8601. Defaults to now. | |
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, so the description does not need to repeat that. It adds useful behavioral context: defaulting to the last 7 days, per-Monitor and per-day aggregates, times in seconds, and a performance score range of 0-100. It could disclose more about edge cases or output shape, but the added context is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three focused sentences. The first front-loads the core purpose and key output semantics; the second routes to alternatives; the third is a concise prerequisite for interpretation. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only summary tool with no output schema, the description covers what it returns (aggregates, trends, units, score range), the defaults, related drill-down tools, and a required prerequisite reading step. This is enough for an agent to select and invoke the tool correctly without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers since and until with ISO-8601 descriptions, and the tool description adds period defaults. projectId lacks a schema description but is clearly implied by 'for a project'. The description also frames since/until meaningfully as the period for the aggregate health picture, which helps the agent understand the parameters' role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific verb ('Get'), resource ('Monitoring Summary'), and scope ('for a project: the aggregate health picture across its Monitors over a period'). It differentiates this aggregate tool from the more granular siblings by explicitly naming list_incidents and list_monitor_cycles as drill-down alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear direction: use this tool for an aggregate monitoring health picture, and 'Drill down instead of asking for more' with specific sibling tools. It also instructs reading the 'monitoring' topic before interpreting results, which helps the agent use the output correctly. This is explicit usage guidance beyond what the schema provides.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_play_statusGet Play StatusARead-onlyIdempotentInspect
Get the current state of a play, with per-step results and log that fill in as it progresses, optionally waiting for it to finish. After play_script, call this with waitSeconds instead of sleeping between polls; the results are complete once running is false. Each step's has-flags (hasBody, hasLog, and their siblings) say which get_step_detail parts would return something for that step, so drill down only where a flag is true. On a Playwright run the whole runner log belongs to the command step that ran the spec file — hasLog points there. A step's screenshotKeys are in capture order; a Playwright step lists the script's own page.screenshot calls in order, with the runner's final screenshot of the test last.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| scriptRunId | Yes | The script run id returned by play_script. | |
| waitSeconds | No | Optional: wait up to this many seconds (capped at 30) and return the moment the play reaches a terminal state. On expiry it returns the current status rather than an error, so just call again for a longer run. Omit for an immediate read. |
Output Schema
| Name | Required | Description |
|---|---|---|
| log | Yes | |
| steps | Yes | |
| errors | No | |
| status | Yes | |
| running | Yes | |
| startedAt | No | |
| finishedAt | No | |
| scriptRunId | Yes | |
| dashboardUrl | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this tool read-only and idempotent, and the description meaningfully supplements that by explaining progressive filling of results, the wait-expiry return behavior, has-flag semantics, Playwright runner log placement, and screenshot ordering. This is rich behavioral context beyond what the annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence contributes operational guidance: core purpose, polling advice, completion condition, drill-down guidance, and output interpretation details. It is structured with the most important usage instruction early, followed by progressively deeper behavioral notes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema and read-only annotations, the description covers everything an agent needs to use it correctly: when to call it, how long to wait, how to detect completion, how to interpret has-flags, and how to understand screenshot and log placement. No critical operational context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with scriptRunId and waitSeconds already documented in the schema. The tool description adds practical context around waitSeconds usage, but projectId remains undocumented and the description does not add meaning for that parameter. Overall, the schema carries most of the parameter documentation burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (a play) and the action (get current state), and further specifies it returns per-step results and logs. It differentiates itself from related tools by referencing play_script and get_step_detail, so an agent can select it correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use this tool: after play_script, call this with waitSeconds instead of polling. It also states when results are complete (running is false) and advises drilling down with get_step_detail only where has-flags are true.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scenarioGet ScenarioARead-onlyIdempotentInspect
Get a single scenario including its Populations, plus a dashboard link and its launch readiness — an empty launchBlockers means the user can launch it as it stands. A field left out of a population is at its default: iterations and iterationsPerUser unlimited, aggressionMultiplier 1.0, and bandwidth unthrottled. Read get_documentation topic 'load-test-scenarios' for what the stages and ramps actually do.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| scenarioId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| name | Yes | |
| maxBots | Yes | |
| version | No | |
| projectId | Yes | |
| fuelBalance | Yes | |
| populations | Yes | |
| dashboardUrl | Yes | |
| estimatedFuel | Yes | |
| totalDuration | Yes | |
| launchBlockers | Yes | |
| populationCount | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds valuable behavior beyond that: it explains what an empty launchBlockers means, and documents defaults for omitted population fields such as unlimited iterations, aggressionMultiplier 1.0, and unthrottled bandwidth. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with front-loaded purpose and then the most important interpretive details. Every sentence earns its place: core resource, launch readiness meaning, population defaults, and a pointer to deeper documentation. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema is present, so return-value structure does not need to be spelled out. The description covers the essential semantics an agent needs: launch readiness, default field behavior, and where to learn about stages/ramps. Nothing critical is missing for correct invocation and interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for parameter meaning. It does not explicitly describe projectId or scenarioId, but these are self-identifying IDs and the description establishes scenario as the resource. Adequate, though not rich.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('a single scenario'), and lists the meaningful contents returned: Populations, a dashboard link, and launch readiness. The 'single' qualifier clearly distinguishes it from sibling tools like list_scenarios.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: retrieving one scenario with its populations, dashboard link, and launch readiness. It does not explicitly name alternatives or exclusions, but the scope is specific enough that an agent can select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_screenshotGet ScreenshotARead-onlyIdempotentInspect
Fetch a screenshot captured during a play as a viewable image. Use the screenshotKey from a step in get_play_status.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| scriptRunId | Yes | ||
| screenshotKey | Yes | The screenshotKey reported by a step in get_play_status. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds the return-type trait ('viewable image') and the source of the key parameter, which is genuine context beyond the annotation hints. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each earning its place: the first states action and output type, the second states the data-source flow. Purpose is front-loaded and there is zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter fetch with rich annotations, the description covers purpose, output format, and the input-flow. The remaining gaps — the precise delivery format of the image and the meaning of the two ID parameters — would improve it but are not critical for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate. It meaningfully explains screenshotKey's provenance (from a step in get_play_status), but projectId and scriptRunId receive no documentation in either the schema or the description; their names are self-explanatory, which prevents a lower score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch') and resource ('a screenshot captured during a play') and clarifies the output is a viewable image. The play-context framing distinguishes it from the many other get_* siblings, none of which deal with screenshots.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Sentence 2 gives clear when-to-use context: this tool consumes a screenshotKey produced by get_play_status, establishing the prerequisite call flow. It stops short of explicitly naming alternatives or exclusions, but no sibling offers a competing screenshot-fetching path, so the omission is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scriptGet ScriptARead-onlyIdempotentInspect
Get a single script including its full command list and variables, plus a dashboard link. The variables are the script's dataset bindings — the ${...} references in its commands resolve through them.
| Name | Required | Description | Default |
|---|---|---|---|
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the operational profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false), lowering the burden. The description adds value beyond annotations by disclosing what the response contains and clarifying the domain semantics of 'variables' as dataset bindings that resolve ${...} references. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler: the first front-loads purpose and return contents, and the second earns its place by disambiguating the potentially confusing term 'variables.' Nothing extraneous is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only getter with two required string parameters, the description covers the essential facts an agent needs: what is fetched, what the response includes, and what 'variables' means. The residual gaps — sibling differentiation and parameter format details — are accounted for in other dimensions, and annotations carry the operational traits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter documentation, but it says nothing about projectId or scriptId beyond what the parameter names imply. The names are self-explanatory, yet the description supplies no format, origin, or relationship guidance for either parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Get a single script') and itemizes the return contents: 'full command list and variables, plus a dashboard link.' The qualifier 'single' distinguishes it from list_scripts, and the content details separate it from sibling getters like get_script_revision and get_script_asset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is provided, and no alternatives are named despite lookalike siblings such as get_script_revision and get_script_asset. The description states what the tool does but leaves the agent to infer when it is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_script_assetGet Script AssetBRead-onlyIdempotentInspect
Read a script asset's content as base64.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | The asset path. | |
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds the base64 encoding detail, which is useful, but does not disclose other behaviors such as error cases or whether path must be relative to a specific root.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one efficient sentence, front-loads the action, and contains no redundant words. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read operation, the description plus annotations provide a reasonable baseline, and the base64 return format is stated. However, with no output schema and sparse parameter documentation, the description leaves some ambiguity about asset path semantics and expected response structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with only 'path' having a generic description. The tool description adds no explanation of projectId, scriptId, or how they relate to identifying a script asset, so the description fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a clear resource ('script asset's content'), and the output encoding ('as base64'). It is clear but does not explicitly distinguish itself from sibling tools like get_script or list_script_assets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when the agent needs to retrieve a script asset's content, and the read-only annotation reinforces this. However, it does not explicitly mention when to prefer this over siblings like list_script_assets or get_script.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scripting_apiGet Scripting APIARead-onlyIdempotentInspect
List the scripting API scopes (bot, http, browser, ...) available in code blocks, or fetch one scope's members with their signatures and docs.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | The scope name, e.g. 'bot' or 'http'. Omit to list the available scopes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds meaningful behavioral detail beyond that by explaining the two modes: omitting scope lists available scopes, while providing one fetches members with signatures and docs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence with no filler. The main action and resource are front-loaded, and the optional behavior is stated compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a low-complexity tool with one optional parameter and no output schema. The description covers both invocation modes and what the agent should expect in response (list of scopes or members with signatures/docs), which is sufficient for correct use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully documents the single optional 'scope' parameter with an example and the omit-to-list behavior, so schema coverage is 100%. The description reinforces the mode distinction but does not add significant new parameter-level detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('List', 'fetch') and names the exact resource: scripting API scopes and their members. The phrase 'available in code blocks' distinguishes this from schema/documentation siblings like get_command_schema or get_documentation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: when an agent needs scripting API scopes or member signatures/docs in code blocks. It does not name alternatives or exclusions, but the tool's purpose is specific enough that the usage context is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_script_revisionGet Script RevisionARead-onlyIdempotentInspect
Get a specific past revision of a script, including its commands, so you can inspect or diff it.
| Name | Required | Description | Default |
|---|---|---|---|
| version | Yes | The revision version number. | |
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose read-only, idempotent, and non-destructive behavior. The description adds that the response includes commands, which is useful beyond the structured fields. However, it does not describe the full return shape or behavior for invalid versions, leaving some ambiguity that annotations do not fully cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every phrase adds meaning: the action, the target resource, the included content, and the intended use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the basic what and why, and annotations cover safety. However, with no output schema, it does not fully describe the return value beyond 'commands,' and it omits guidance that valid version numbers might come from list_script_revisions. This leaves moderate gaps for an agent trying to call it correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with only 'version' described. The tool description does not compensate by explaining 'projectId' or 'scriptId' beyond their self-evident names, and it does not clarify required formats or relationships between the parameters. This is a significant gap given the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Get a specific past revision of a script.' It also adds 'including its commands,' which distinguishes it from list_script_revisions and get_script by clarifying that it returns executable content, not just metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'So you can inspect or diff it' provides a clear use case for when this tool is appropriate. It does not explicitly name alternatives or exclusions like get_script or restore_script_revision, but the intended context is evident from the wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_step_detailGet Step DetailARead-onlyIdempotentInspect
Read the payloads get_play_status leaves out — response bodies, headers, captured values, console messages, network activity, page HTML, logs — for one or more steps of a play, choosing the parts per call. Ask for several steps at once: a failing request and the earlier step that captured the token it used is one call. Use this instead of adding a validator or code block to the script just to echo a response back to yourself. Parts are bot-type-shaped, and asking for a part a run has no source for returns an error naming what it does have.
| Name | Required | Description | Default |
|---|---|---|---|
| parts | Yes | The parts to return: body, headers, captured, console, network, network_full, html, or log. Required, so a drill-down never returns more than you asked for. 'network' ranks the resources (failures, non-2xx, and the slowest 20) with the totals beside them; 'network_full' returns every one. | |
| selector | No | Optional CSS selector, for the 'html' part only. Without it, html returns an outline of the page's anchorable elements with a suggested selector for each — ask for that first when you don't know what to target. With it, html returns the matching elements' outerHTML. | |
| projectId | Yes | ||
| scriptRunId | Yes | The script run id returned by play_script. | |
| stepIndexes | Yes | The step indexes to drill into, starting at 1, as reported by get_play_status. |
Output Schema
| Name | Required | Description |
|---|---|---|
| steps | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful behavior beyond that, such as the ability to request multiple steps in one call, the bot-type-shaped nature of parts, and the error behavior when requesting a part with no source available.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a focused four-sentence paragraph with the main purpose front-loaded and no redundant filler. The phrase 'bot-type-shaped' is somewhat cryptic and could obscure the meaning, but overall the structure is efficient and every sentence contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the annotations, and the presence of an output schema, the description covers the core purpose, common usage patterns, and important error behavior. It does not fully explain 'bot-type-shaped' parts or exhaustively cover all parameters, but the schema fills those gaps and the description is sufficient for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so most parameters are already documented. The description adds meaning by explaining the parts payload conceptually and by illustrating that stepIndexes can combine dependent steps, such as a failing request and the step that captured a token, in a single call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action and resource: it reads the payloads get_play_status leaves out, such as response bodies, headers, captured values, console messages, network activity, HTML, and logs, for one or more steps of a play. It also explicitly distinguishes itself from get_play_status by describing what this tool adds beyond that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when you need detail beyond get_play_status, and it explicitly recommends using this instead of adding a validator or code block to echo a response. It does not explicitly state when to prefer get_play_status or other alternatives, so the guidance is strong but not fully exclusionary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_variable_schemaGet Variable SchemaARead-onlyIdempotentInspect
Get the schema for binding a script variable to a dataset, including the accepted update and selection strategies. A variable is a named row cursor bound to one whole dataset by dataSetId: on each pull it holds an entire row, advancing per occurrence, per iteration, or per bot ('vuser') depending on updateStrategy, with rows taken sequentially or randomly per selectionStrategy. Reference it as ${name} in command fields or bot.getVariable('name') in Playwright code. After changing bindings or references, run validate_script to cross-check them.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the call read-only/idempotent/non-destructive, and the description adds substantial behavioral context: variables act as cursors, advancing per occurrence/iteration/bot depending on updateStrategy and choosing rows sequentially/randomly via selectionStrategy. It also documents reference syntax and the follow-up validation step, going well beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: purpose, variable semantics, reference syntax, and post-edit validation. The most important identification information is front-loaded and the description avoids repetition of the title or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument metadata getter, the description covers what the schema contains and how to use the result, and it works well with the read-only annotations. It could be more complete by explicitly describing the shape/format of the returned schema, but the lack of an output schema makes this a minor gap rather than a critical omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is no missing parameter documentation; the baseline for no-parameter tools is 4. The description adds useful domain semantics (dataSetId, updateStrategy, selectionStrategy) even though those are concepts in the returned schema rather than call parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb and resource: 'Get the schema for binding a script variable to a dataset,' including update/selection strategies. This is clearly distinct from sibling tools like get_command_schema because it focuses on script variable-to-dataset binding.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the tool is relevant: when working with script variable bindings and their strategies, and it advises running validate_script after changing bindings/references. It does not explicitly name alternatives or state when not to use it, but the intended use is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_scriptImport ScriptADestructiveInspect
Import an open-source load test (JMeter .jmx, k6 .js, or Locust .py) into an existing script, converting it to Loadster commands. An import may create script assets, datasets, and a scenario; replace=true also replaces the script's existing commands and variables. Returns action items to review. For OpenAPI/Swagger or Postman, read the spec yourself and build commands with create_script instead.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | The source tool: jmeter, k6, or locust. | |
| baseUrl | No | Optional base URL to prefix relative requests. | |
| content | Yes | The source file content (the .jmx XML, .js, or .py text). | |
| replace | No | Replace existing commands (true) or append to them (false). Defaults to false. | |
| scriptId | Yes | The script id to import into (create an empty script first). | |
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and readOnlyHint=false, and the description elaborates with critical side effects: it may create script assets, datasets, and a scenario, and replace=true replaces existing commands and variables. This adds meaningful behavioral detail that annotations alone do not convey, such as the scope of destruction and what gets created. It does not contradict annotations, and the description's mention of 'replace' aligns with destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each earning its place: the first defines the action and inputs, the second summarizes side effects, and the third provides an exclusion with a clear alternative. It is front-loaded with the core purpose and uses no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (parsing multiple formats, side effects, replace semantics) and the absence of an output schema, the description is thorough: it covers what inputs are needed, what happens on success (returns action items), the meaning of replace, and the exclusion case. No critical operational detail is missing for an agent to decide whether to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, leaving the projectId parameter undocumented in the schema. The description does not explicitly describe each parameter, but it explains the overall purpose of 'source', 'content', and 'replace', and it mentions that import may create assets/datasets/scenario, which gives context for scriptId and projectId. While it doesn't fully compensate for the missing projectId description, the schema covers most parameters, so a slight bonus above baseline is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (import) and resource (open-source load test scripts into Loadster), enumerates the supported source formats, and explicitly contrasts with OpenAPI/Swagger and Postman, which are not supported. This clearly differentiates the tool from siblings like create_script and makes its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool (for JMeter, k6, Locust files) and when not to (for OpenAPI/Swagger or Postman), even naming the alternative tool (create_script) to use instead. This goes beyond basic context and gives actionable routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_command_typesList Command TypesARead-onlyIdempotentInspect
List the command (step) types available for a bot type, with a description and the bot types each applies to. Use this before authoring commands.
| Name | Required | Description | Default |
|---|---|---|---|
| botType | No | The bot type: HTTP, BROWSER, or PLAYWRIGHT. Omit to list all types. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context about what results contain (description and bot types), but no further behavioral traits such as pagination, ordering, or output format are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The core action and output are front-loaded, and the optional usage hint is concise and actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional enum parameter, no output schema, and strong read-only annotations, the description fully covers what the tool does, what it returns, and when to use it. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single optional botType parameter is fully documented with an enum and the 'Omit to list all types' behavior. The description adds no additional parameter semantics beyond the schema, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('List') and resource ('command (step) types') scoped by bot type, and notes the returned information includes descriptions and applicable bot types. It does not explicitly distinguish itself from sibling tools like get_command_schema, but the listing purpose is evident and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context with 'Use this before authoring commands,' which tells the agent when the tool is appropriate. It does not mention alternatives or when-not conditions, but the guidance is clear enough for a simple listing tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_datasetsList DatasetsARead-onlyIdempotentInspect
List the datasets in a project with their id, name, and row/column counts (without the full value table).
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds meaningful behavioral context by specifying exactly what is returned and, importantly, what is excluded (the full value table), which helps set agent expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that leads with the action and resource, then packs essential return-field details and an exclusion into minimal words. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter, read-only list operation, the description is complete: it states the resource, scope, returned fields, and what is deliberately omitted. Annotations cover the safety profile, and no output schema exists to explain return values, so the description carries the burden well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter, projectId, and the description references it via 'in a project', giving some context. However, schema description coverage is 0%, and the description does not elaborate on the expected format, semantics, or error behavior of projectId beyond what the parameter name already implies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List', the resource 'datasets', the scope 'in a project', and the specific returned fields (id, name, row/column counts). It also explicitly distinguishes itself by noting it does not return the full value table, which separates it from get_dataset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this tool enumerates dataset metadata and sizes without pulling full data. It does not explicitly name alternative tools or say when not to use it, but the parenthetical '(without the full value table)' implicitly steers the agent away from using this when dataset contents are needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_enginesList EnginesARead-onlyIdempotentInspect
List everywhere a population can run, with the customer's own limits a scenario is written against. Each entry's id is what a population's loadEngineId names — never invent one. Type CLOUD is a single Loadster cloud region, WILDCARD is a pattern ('north-america-', '') that lets Loadster pick a region inside it when the test launches, and PRIVATE is one of the customer's own Loadster Engines, the only kind that can reach a host on their network. The per-population caps and stage bounds are enforced when a scenario is saved; capacities, locked regions, and the per-region and per-engine caps matter only at launch and never stop a save. While limits.freeTrialCustomer is true, an entry with locked true can be saved onto but not launched from, so put a trial's populations on an entry with locked false and keep the scenario's peak within limits.maxFreeTrialBotsPerTest.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description discloses nontrivial behavior: which limits are enforced at save time versus launch time, locked entries being save-but-not-launch for free trials, and the meaning of WILDCARD matching at launch. This is exactly the kind of context that prevents misuse and exceeds what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded: the core purpose is in the first sentence, and every subsequent sentence adds a distinct piece of operational knowledge. There is no filler or repetition; each sentence earns its place given the subtle engine semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining what the entries are and how to interpret them, and it does so thoroughly. It covers id provenance, type semantics, save-time vs launch-time limits, and the free-trial caveat, so an agent has enough to invoke the tool and consume its result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema coverage is effectively complete, so there are no input parameters to document. The description still adds value by explaining the semantic meaning of output fields (id, type, locked, limits). Baseline 4 for a zero-parameter tool is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific action and resource: list all places a population can run, i.e. the customer's engines/regions. It goes on to distinguish CLOUD, WILDCARD, and PRIVATE entries, making clear this is the engine/load-engine listing tool rather than a generic list. This fully differentiates it from sibling list_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the intended use clear: consult this tool to get authoritative loadEngineId values and never invent an engine id. It also provides a concrete decision rule for trial customers (use locked false entries and stay within maxFreeTrialBotsPerTest). It does not explicitly name alternative tools or say when not to use it, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_incidentsList IncidentsARead-onlyIdempotentInspect
List a project's Incidents — open ones first, then the most recently closed; unresolved only by default, with includeResolved for incident archaeology over a since/until window and monitorId to ask about one Monitor. An Incident opens when a Monitor's consecutive failing Cycles reach its failure threshold, and closes after enough consecutive passes, or when the Monitor is disabled. Rows are lean; read one Incident's notifications, statuses, and cycle references with get_incident. The 'monitoring' topic covers how Incidents notify.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many Incidents to return (default 10, maximum 50). | |
| since | No | Only Incidents still active after this ISO-8601 instant (2026-08-26T00:00:00Z) or date (2026-08-26). | |
| until | No | Only Incidents opened at or before this ISO-8601 instant or date. | |
| monitorId | No | Only this Monitor's Incidents, as reported by list_monitors. | |
| projectId | Yes | ||
| includeResolved | No | Include resolved Incidents too. Defaults to false: only open Incidents. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description adds sorting order, default unresolved-only behavior, Incident lifecycle semantics, and the 'lean rows' expectation. This is meaningful behavioral context that helps an agent predict results and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and mostly front-loaded with the core behavior. The lifecycle explanation is useful, though phrases like 'incident archaeology' and the final monitoring-topic sentence add color rather than critical operational guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers sorting, default filtering, optional filters, lifecycle behavior, and directs users to get_incident for detailed fields. With no output schema, it does not enumerate exact returned fields, but 'lean rows' and the pointer to get_incident provide adequate guidance for agent selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is high (83%), so the description need not restate parameter meanings. It adds some context around since/until as a window and includeResolved as 'incident archaeology', but most parameter semantics are already in the schema. The projectId parameter lacks a schema description, but its meaning is obvious from the tool name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists a project's Incidents with specific sorting and default filtering behavior. It also distinguishes itself from get_incident by noting that list rows are lean and details are available from get_incident.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly explains when to use includeResolved, since/until, and monitorId, and directs users to get_incident when they need richer detail on a single Incident. This gives an agent clear decision criteria among sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_load_testsList Load TestsARead-onlyIdempotentInspect
List a project's finished Load Tests, newest first, a window of summary rows at a time. A Load Test is a run of a Scenario at scale by many bots; the user launches one from the Dashboard, it appears here once it finishes, and nothing on this surface launches, stops, or deletes one. Read one test's Test Report with get_load_test_report.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many tests to return (default 10, maximum 50). | |
| offset | No | Zero-based index of the first test to return (default 0). | |
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description adds meaningful behavior beyond that: it returns only finished tests, orders newest first, returns summary rows in windows, and explicitly states that the surface does not launch, stop, or delete tests. This is consistent with the annotations and adds useful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then provides a concise domain definition and a pointer to the related report tool. Every sentence earns its place, and there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list operation, the description covers the key facts: what is listed, the ordering, the pagination style, the lifecycle context, and how to get more detail on a single test. The lack of an output schema is mitigated by the phrase 'summary rows,' and annotations cover the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers limit and offset with descriptions, but projectId has no description. The description implies projectId is the project identifier by saying 'a project's finished Load Tests,' but it does not add detailed parameter semantics. With 67% schema coverage, the description provides only marginal additional meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'List a project's finished Load Tests, newest first, a window of summary rows at a time.' It clearly distinguishes itself from get_load_test_report by stating that reading one test's Test Report is a separate action, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: use this to list finished Load Tests and use get_load_test_report to read a single test's report. It also clarifies what this surface does not do (launch, stop, or delete tests), which helps avoid misuse, though it does not exhaustively enumerate all alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_monitor_cyclesList Monitor CyclesARead-onlyIdempotentInspect
List a Monitor's Cycles, newest first — one row per Cycle with its status, timestamps, the engine id it ran from, and headline response times; a Cycle with no status yet is scheduled or still running. Window with since/until and limit to spot patterns — flapping around a threshold, slow drift, failures from one region — then drill into one Cycle's full metrics and failure message with get_monitor_cycle_detail. The 'monitoring' topic explains what each status means.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many Cycles to return (default 20, maximum 100). | |
| since | No | Only Cycles scheduled at or after this ISO-8601 instant (2026-08-26T00:00:00Z) or date (2026-08-26). | |
| until | No | Only Cycles scheduled before this ISO-8601 instant or date (exclusive). | |
| monitorId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only, idempotent, and non-destructive, so the bar is lower. The description adds valuable behavior: newest-first ordering, one row per Cycle, included fields, and the meaning of a Cycle with no status yet. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no wasted words. It front-loads the core behavior and output, then adds usage context and a pointer to related documentation, all in an efficient structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining the return shape; it does so clearly with fields, ordering, and status semantics. It also gives filtering guidance, an alternative tool, and a doc pointer, making it complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents limit, since, and until with formats and constraints. The description adds contextual value by suggesting since/until/limit for pattern spotting, but it does not add meaning for monitorId or projectId beyond what their names imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List a Monitor's Cycles'), specifies ordering ('newest first'), and describes the row-level fields returned. It clearly distinguishes this from the sibling get_monitor_cycle_detail by framing this as the overview listing tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the exact use case: spotting patterns like flapping, slow drift, or regional failures. It also names the alternative tool for deeper investigation ('drill into one Cycle's full metrics and failure message with get_monitor_cycle_detail'), giving clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_monitoring_locationsList Monitoring LocationsARead-onlyIdempotentInspect
List everywhere a Monitor's Cycles can run from. Each entry's id is what a Monitor's engineIds names — never invent one. Type CLOUD is one of Loadster's shared monitoring locations, and PRIVATE is one of the customer's own Loadster Engines, the only kind that can reach a host on their network. Monitoring locations are their own set, not the load test regions list_engines offers. Read this before create_monitor or update_monitor.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, so the safety profile is covered. The description adds meaningful behavioral context: CLOUD vs PRIVATE semantics, PRIVATE being the only kind that can reach customer-network hosts, and the ID-reuse warning. It could mention output format or pagination, but that's not critical for a 0-parameter read-only list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: purpose, ID warning, CLOUD/PRIVATE explanation, and relationship to siblings. The core purpose is front-loaded and there is no redundant restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only list with rich annotations, this is complete. It tells the agent what the entries mean, how they relate to other tools, and when to consult this endpoint, so no additional context is needed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the schema needs no elaboration. The description compensates by explaining the meaning of the returned entries' types and how the ids are consumed by other tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List everywhere a Monitor's Cycles can run from.' It distinguishes itself from list_engines by explicitly noting monitoring locations are their own set, not the load test regions list_engines offers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: 'Read this before create_monitor or update_monitor.' It also warns against inventing IDs, tells the agent these IDs are what engineIds names, and names the sibling list_engines as offering a different set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_monitorsList MonitorsARead-onlyIdempotentInspect
List the Monitors in a project. A Monitor is a scheduled, low-volume check that runs a Script on a frequency to verify health and alert on failure; each execution is a Cycle. Each row is a summary without the thresholds and engine ids; read one Monitor's full configuration with get_monitor. Read get_documentation topic 'monitoring' for how Monitors, Cycles, and Incidents behave.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is known. The description adds valuable behavioral context beyond annotations: it explains that Monitors are scheduled low-volume checks, that each execution is a Cycle, and that rows are summaries without thresholds/engine ids. It also directs to documentation for lifecycle behavior, which is useful. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no redundancy. The core action is front-loaded, the domain context is concise, and the pointer to get_monitor and documentation is placed at the end. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list operation with one parameter and no output schema, the description is nearly complete: it covers scope, summary-level result, and how to get more detail. It does not describe pagination, sorting, or the exact fields returned, but those are typically minor for such a tool and the pointer to get_monitor fills the main gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter projectId is obvious from the description ('in a project') and the required field name is self-explanatory. However, schema description coverage is 0%, and the description does not specify the format or constraints of projectId (e.g., is it a UUID or name?). Since there is only one parameter and its role is clear, the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the Monitors in a project'), defines the domain concept (Monitor/Script/Cycle), and clearly distinguishes itself from get_monitor by noting each row is a summary. This differentiates it from sibling tools like list_monitor_cycles and get_monitor without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when you need a summary list of Monitors in a project, and explicitly points to get_monitor for full configuration. However, it does not explicitly state when not to use it (e.g., for cycles or incidents), though the context makes this reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsList ProjectsARead-onlyIdempotentInspect
List the Loadster projects the current token can access, each with its id and name. Every other tool takes one of these ids. An empty list means the connected team has no projects yet: nothing on this surface creates one, so ask the user to create a project in the Dashboard, then call this again.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description adds meaningful behavioral context: results are scoped to the current token, each result includes id and name, and an empty list has a specific meaning. It also correctly notes that no operation on this surface creates projects, which is useful beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler. The core purpose is front-loaded, the cross-tool dependency is stated in the second sentence, and the empty-list handling is a compact, valuable addition. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless list operation with no output schema, the description is complete: it covers access scope, result fields, empty-list meaning, and the correct user-facing action for that case. No important detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema coverage is 100%, so there is no parameter meaning for the description to add. The description focuses instead on output semantics, which is appropriate for a parameterless tool and meets the baseline for this case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a specific resource ('Loadster projects'), and the access scope ('current token can access'). It also names the returned fields (id and name), making the tool's purpose unmistakable and distinguishable from the many other list_* sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly positions this tool as the prerequisite for every other tool by stating 'Every other tool takes one of these ids.' It also gives concrete guidance for the empty-list edge case: nothing on this surface creates a project, so the user should create one in the Dashboard and call this again. This is actionable and clarifies when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_scenariosList ScenariosARead-onlyIdempotentInspect
List the scenarios in a project with their id, name, version, population count, peak bots, and total duration in milliseconds — without the populations, fuel estimate, and launch blockers get_scenario reports. A Scenario is a saved load test configuration: a named list of Populations (the Dashboard calls them "bot groups"), each of which runs one script with its own load shape on one engine or region. Saving a scenario costs nothing and runs nothing; a user launches it from the Dashboard.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds useful behavioral context: it returns summary data only and omits heavy fields like populations and launch blockers, and it clarifies that scenarios are saved configurations that are not executed by saving.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads the core action and output fields, and the subsequent sentences provide useful domain context about scenarios and their cost model. Every sentence contributes, though the scenario definition could be seen as slightly more detail than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one parameter and no output schema, the description is largely complete: it documents the resource, the fields returned, the relationship to get_scenario, and relevant scenario semantics. It does not discuss pagination, ordering, authorization, or error cases, but these are not critical for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only provides a required projectId string with no description, so the description carries the burden. It clarifies that the operation lists scenarios 'in a project', giving meaningful context to projectId. The description also explains scenario and population concepts, although it does not specify ID formats or error behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List the scenarios in a project') and enumerates the exact fields returned. It also explicitly distinguishes itself from get_scenario by naming what it omits, so an agent can reliably tell sibling tools apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly contrasts list_scenarios with get_scenario by listing which fields each provides, implying when to choose the lighter overview tool. It does not provide an explicit 'use when / do not use when' rule for every alternative, but the differentiation from get_scenario is strong enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_script_assetsList Script AssetsARead-onlyIdempotentInspect
List the files attached to a script (request bodies, upload files, Playwright support files). Each path is the key a command references.
| Name | Required | Description | Default |
|---|---|---|---|
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safe read-only profile is covered. The description adds useful behavioral context by defining the asset types and explaining how paths are used as command reference keys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The main action and scope are stated immediately, and the path-as-key clarification is meaningful rather than redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list operation with rich annotations and only two straightforward parameters, the description is nearly complete. It could explicitly state the return shape or empty behavior, but the path-as-key sentence provides enough functional context, especially since no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter documentation. It never explains scriptId or projectId, how they relate, or where they come from. The parameter names are self-explanatory, but the description adds no direct parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'List', and clearly identifies the resource: files attached to a script, with concrete examples (request bodies, upload files, Playwright support files). This distinguishes it from siblings like list_scripts or get_script_asset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when you need to enumerate the files attached to a script and adds that each path is the key a command references. However, it does not explicitly state when to prefer this over get_script_asset or put_script_asset, nor does it mention any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_script_revisionsList Script RevisionsARead-onlyIdempotentInspect
List a script's past revisions (newest first) with version, save time, author, and step count.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | How many revisions to return (default 10, maximum 100). | |
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds useful behavioral context beyond that: results are newest-first and include version, save time, author, and step count. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One clean, front-loaded sentence that states the operation, scope, ordering, and returned fields without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with strong annotations, the description covers enough: result shape and ordering are stated, and count is documented in the schema. It would be more complete with minimal guidance on projectId/scriptId, but names and context make this mostly self-evident.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% and the description does not compensate. scriptId and projectId have no descriptions and are not explained in the tool description; only count is documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb 'List' with resource 'a script's past revisions', and adds ordering ('newest first') and result contents. This clearly distinguishes it from singular get_script_revision and restore_script_revision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: call it to view revision history for a script. However, it does not explicitly state when to prefer it over get_script_revision or restore_script_revision, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_scriptsList ScriptsARead-onlyIdempotentInspect
List the scripts in a project with their id, name, bot type, step count, and current revision version (without the full command list).
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, and non-destructive behavior. The description adds meaningful behavioral context beyond the annotations by specifying exactly which fields are returned and that the command list is deliberately omitted. This helps set expectations about response content without repeating annotation information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-structured sentence that front-loads the action and resource, lists concrete return fields, and adds the key exclusion at the end. Every part contributes useful information with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, read-only listing tool with one required parameter, the description is complete: it states the resource scope, enumerates return fields, and notes the command-list exclusion. The annotations cover safety and idempotency, so no additional behavioral context is necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the parameter name 'projectId' with no description, and schema coverage is 0%. The description says 'in a project' but does not explicitly explain that the required projectId identifies the project or mention how to obtain a valid projectId. It partially compensates, but the parameter semantics are still somewhat underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('List') and resource ('scripts in a project'), and enumerates the exact returned fields: id, name, bot type, step count, and current revision version. It also differentiates itself by explicitly saying it returns 'without the full command list', distinguishing it from get_script and list_script_revisions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear this is a summary listing tool and that full command lists are not included, implying that a different tool should be used when full commands are needed. It does not explicitly name an alternative sibling, but the usage context is sufficiently clear for an agent to choose it for project-level script summaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_scriptPlay ScriptAInspect
Start playing a saved script. Returns immediately with a run id and a queued status, since a play can take anywhere from a few seconds (HTTP) to several minutes (browser or Playwright); read the results with get_play_status.
| Name | Required | Description | Default |
|---|---|---|---|
| scriptId | Yes | The script id to play. | |
| projectId | Yes | ||
| skipWaitTimes | No | Skip wait/think times for a faster run. Defaults to false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the most important runtime behavior beyond annotations: the call returns immediately with a run id and queued status, and execution may take from seconds to minutes depending on the target. It also tells the agent where to read results. This goes well beyond the readOnly, idempotent, and destructive hints, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence that front-loads the action and proceeds through immediate return, duration context, and follow-up route. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an asynchronous tool with no output schema, the description provides the essential contract: what is returned now (run id, queued status), why the wait may be long, and how to get results. It covers the invocation context and follow-up, making it complete enough for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes scriptId and skipWaitTimes; the description does not elaborate on projectId or add parameter-level constraints beyond the schema. Since coverage is 67% and the remaining parameter name is self-explanatory, the description adds little parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with the specific verb 'Start playing a saved script,' making the action and resource explicit. It also distinguishes itself from siblings by noting the return behavior and pointing to get_play_status, so an agent can tell it apart from create, validate, delete, and list operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes the context: use this to initiate execution of a script that already exists, then poll for results. It explicitly routes to get_play_status for reading results, but it does not list exclusions or when-not conditions versus other script operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
put_script_assetPut Script AssetADestructiveInspect
Upload or replace a script asset from base64 content. Use the path as the key a command references — e.g. set an HTTP command's requestBody to { "key": "" } to give it a request body, or list the path in a browser upload step's files.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | The asset path / key to store it under, e.g. 'body.json'. | |
| scriptId | Yes | ||
| projectId | Yes | ||
| contentBase64 | Yes | The file content, base64-encoded. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the agent knows this is a mutating operation. The description adds the 'upload or replace' behavior, which implies overwriting existing assets at the same path. It doesn't disclose whether replacement is destructive to existing content or if there are any side effects, but the annotations cover the safety profile. The description adds some context about how the asset is used (as a command key) but not deep behavioral details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the core action ('Upload or replace a script asset from base64 content') and then provides a useful example. It's concise and every sentence earns its place, though the example could be seen as slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with destructiveHint=true, the description explains the primary use case and the path-key relationship. However, it doesn't mention what happens on replacement (e.g., overwriting existing content), any prerequisites (e.g., project/script must exist), or the response format. With no output schema, the agent might not know what to expect. The description is adequate but has gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, with path and contentBase64 having descriptions. The description adds meaning by explaining the path's role as a key for commands, which goes beyond the schema's 'The asset path / key to store it under.' However, scriptId and projectId have no descriptions in the schema, and the description doesn't compensate for those. The description adds some value but doesn't fully cover the undocumented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Upload or replace a script asset from base64 content.' It specifies the resource (script asset) and the action (upload/replace), and distinguishes it from siblings like get_script_asset and delete_script_asset by focusing on the write operation. The example of using the path as a key for commands adds concrete context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool: when uploading or replacing a script asset, and explains how the path is used as a key for commands. It doesn't explicitly state when not to use it or name alternatives, but the sibling list includes get_script_asset and delete_script_asset, and the description's focus on upload/replace makes the usage context clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_script_revisionRestore Script RevisionBDestructiveInspect
Restore a script to a past revision, saving that revision's commands as a new revision.
| Name | Required | Description | Default |
|---|---|---|---|
| version | Yes | The revision version number to restore. | |
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag this as non-read-only and destructive, so the description's addition of 'saving that revision's commands as a new revision' adds useful behavioral context about how the restore is implemented. However, it does not clarify what happens to the current revision or draft, potential side effects, or whether the operation is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every phrase earns its place, and the key behavior is communicated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive operation with no output schema and sparse parameter documentation, the description conveys the core behavior but leaves important gaps: parameter meanings, what happens to the current state, and what the agent should expect in return. It is serviceable but not fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only 'version' has a schema-level description; 'projectId' and 'scriptId' are bare string fields. The description does not explain the purpose or relationship of these parameters, and with 33% schema coverage it is not enough for an agent to confidently populate them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Restore a script to a past revision,' and clarifies the operation is recorded as 'a new revision.' This makes the tool's intent clear and distinguishes it from read-only revision tools like get_script_revision and list_script_revisions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as update_script, duplicate_script, or get_script_revision. The context is only implied by the word 'restore,' with no exclusions or conditions for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_scriptStop ScriptAIdempotentInspect
Stop a running play. Errors if the run does not exist; once the stop takes effect, get_play_status reports the run as stopped.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| scriptRunId | Yes | The script run id to stop. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a mutating, non-destructive, idempotent operation. The description adds useful behavior: errors if the run does not exist, and the stop is not necessarily instantaneous ('once the stop takes effect'), plus the post-condition observable through get_play_status. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, then concise behavioral details. Every sentence earns its place without redundancy or unnecessary filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity two-parameter tool with no output schema, the description covers the core action, error behavior, and post-condition, and even points to a verification tool. The main gap is the undocumented projectId semantics, but overall the agent can select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: scriptRunId is documented, but projectId has no description and the tool description does not explain it. The description adds no meaning for the required projectId parameter, leaving an important invocation detail to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Stop a running play.' It also clarifies the error condition and expected observable outcome, referencing get_play_status, which distinguishes it from related run-management tools like play_script and get_play_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: use this tool to stop a running play, and verify the outcome via get_play_status. It implicitly contrasts with play_script by using 'stop,' though it does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_feedbackSubmit FeedbackAInspect
Optional. Sends a note about this tool surface to the Loadster team, who read it to fix what the surface got wrong. Nothing requires it and no task depends on it. A note is about the tools themselves: what worked, what was missing or confusing, what needed a workaround. It should not carry conversation data: the user's prompts, their credentials, or anything from their account. Feedback is at most 2000 characters; a longer submission is rejected rather than truncated.
| Name | Required | Description | Default |
|---|---|---|---|
| outcome | Yes | How the task ended. | |
| feedback | Yes | What worked and what did not about the tools, in at most 2000 characters. | |
| taskDescription | No | The kind of task the note is about, in your own summary rather than the user's wording, in at most 200 characters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While annotations are sparse (readOnlyHint=false, etc.), the description discloses critical behavior: that feedback is optional, has a 2000-character limit, that longer submissions are rejected rather than truncated, and that it must not carry sensitive data. It clearly indicates that this is a non-essential side effect, which is valuable beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the 'Optional' tag and purpose, and every sentence earns its place. It covers what the tool does, when to use it, content constraints, and length behavior without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that the tool is a simple side-action with no output schema, the description is nearly complete. It covers the purpose, optional usage, content restrictions, and rejection on length. However, it does not mention what happens on successful submission (e.g., response or confirmation), but given the low complexity and lack of output schema, this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, meaning the schema already documents each parameter (outcome, feedback, taskDescription). The description adds context about the feedback content and length limit, but does not significantly add detail beyond what the schema provides. Baseline 3 is appropriate because the schema does the heavy lifting, and the description complements with usage context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this tool sends feedback about the tool surface to the Loadster team)Skip. It explicitly notes that nothing requires it and no task depends on it, which disambiguates its purpose from other action tools. The verb is specific ('sends a note') and the resource is clear ('about this tool surface').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says 'Optional' and explicitly states when it is not needed ('Nothing requires it and no task depends on it'). It clarifies feedback content boundaries: what it should and should not contain (excluding conversation data) and warns about the length limit being rejected. This guidance is direct and excludes misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_datasetUpdate DatasetADestructiveInspect
Replace an existing dataset's name and values — the values given become the entire table, so send the complete table including every row you want to keep. To add rows, call append_dataset_rows instead.
| Name | Required | Description | Default |
|---|---|---|---|
| dataset | Yes | The new dataset name and table of values (data rows only — no header row). | |
| datasetId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, and the description adds valuable behavioral context beyond those flags: the values given become the entire table, so the caller must send the complete table. This clarifies the exact nature of the destructive replacement, which is useful for an agent making the call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core action and the critical data-completeness warning are front-loaded, and the sibling alternative is stated concisely. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the operation's destructive replacement behavior and the key alternative tool. With no output schema, return-value explanation is not required. It doesn't mention prerequisites like the dataset needing to exist, but that is reasonably implied by 'Replace an existing dataset'. Overall, it is complete enough for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only dataset is described in the schema). The description compensates by clarifying the dataset parameter's role: it carries the new name and the full table of values. It also reinforces that values are data rows only. However, projectId and datasetId are not given additional meaning, so the compensation is incomplete but still strong.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Replace an existing dataset's name and values'. It also explicitly distinguishes itself from append_dataset_rows by telling the agent to call that sibling when adding rows, so the purpose is unambiguous and well differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the condition for using this tool ('replace an existing dataset's name and values') and provides a clear alternative for a different case ('To add rows, call append_dataset_rows instead'). This gives the agent direct routing guidance without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_load_test_notesUpdate Load Test NotesADestructiveIdempotentInspect
Replace the notes on a finished Load Test's Test Report — the only writable thing on a Load Test. Ask the user before writing: notes belong to the user, and advising on a test never requires writing them. The text sent becomes the report's entire notes, so first read the current notes with get_load_test_report and preserve the user's content in what you send.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | Yes | The full notes text to store on the Test Report, replacing whatever notes it already carries. At most 4096 characters; an empty string clears the notes. | |
| testId | Yes | The load test id, as reported by list_load_tests. | |
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description matches the annotations: destructiveHint=true and readOnlyHint=false, and idempotentHint=true is consistent with a replace operation. Beyond the annotations, it discloses that the tool is the only writable field on a Load Test, that it overwrites the entire notes, and that an empty string clears notes. Minor gap: it doesn't mention potential effects on other users or restore/versioning, but it adds substantial context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loads the core behavior, and every sentence earns its place. It covers the action, the guardrail, and the overwrite caveat without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and three required parameters, the description addresses the critical operational context: when to write, what to preserve, and the full-replacement behavior. The sibling list and annotations cover the rest, so an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the description adds meaningful semantics beyond the schema: it emphasizes that notes replaces the entire report notes, advises reading current notes first to preserve content, and clarifies the user-ownership aspect. This goes beyond the schema's basic parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific verb and resource: it replaces the notes on a Load Test's Test Report, emphasizing that it is the only writable thing on a Load Test. This distinguishes it from all sibling update tools and get_load_test_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: ask the user before writing, never write when merely advising on a test. It also directs the agent to first read the current notes with get_load_test_report and preserve the user's content, and warns that the text sent becomes the entire notes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_monitorUpdate MonitorADestructiveInspect
Replace an existing Monitor's exposed configuration: name, Script, frequency, engine ids, failure and recovery thresholds, tags, and the acceptable-metric thresholds. An omitted threshold is cleared and omitted tags are removed, so read the current configuration with get_monitor first and send everything that should remain. Everything outside those fields is untouched: an update never changes whether the Monitor is enabled, and never touches its notification policies or maintenance windows, so a config edit cannot stop a running Monitor, start a stopped one, or unhook the user's alerting.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The new monitor name. | |
| tags | No | Tag names to organize the Monitor. Prefer the names other Monitors already carry, as list_monitors reports them — a name that doesn't exist yet becomes a new tag, so don't invent tags the user didn't ask for. Omit for no tags. | |
| scriptId | Yes | The id of the Script in this project the Monitor plays each Cycle. | |
| engineIds | Yes | The ids of the monitoring locations the Cycles run from, as list_monitoring_locations reports them; at least one is required. Monitoring locations are their own set, not the load test regions list_engines offers, and an unknown id is refused with an error naming the valid ones. | |
| frequency | Yes | How often a Cycle runs, in milliseconds: 60000 (one minute) to 86400000 (one day). | |
| monitorId | Yes | ||
| projectId | Yes | ||
| thresholds | No | The acceptable-metric limits a Cycle is judged against, each optional within the object: responseTimeAverageMin/Max, responseTimeTotalMin/Max, cycleDurationMin/Max, timeToFirstByteMax, firstContentfulPaintMax, largestContentfulPaintMax, and totalBlockingTimeMax in milliseconds; cumulativeLayoutShiftMax as the unitless Web Vitals score; performanceScoreMin as the 0-100 floor. A limit left out is not checked; omit the whole object to check none. | |
| failureThreshold | No | Consecutive failed Cycles before an Incident opens (1 to 128; default 1). | |
| recoveryThreshold | No | Consecutive passing Cycles before an open Incident closes (1 to 128; default 1). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description discloses the replace semantics: omitted thresholds are cleared, omitted tags are removed, and everything outside the listed fields is untouched. This gives the agent a precise model of side effects and prevents surprise data loss.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the action and field list, then the critical clearing behavior, then the non-scope safety boundary. No filler, no repetition of the title, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutating tool with nested objects and no output schema, the description is complete: it identifies what is replaced, what is cleared if omitted, what to do before calling, and what is guaranteed not to change. An agent has enough information to decide when to call it and how to construct the request.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even with 80% schema coverage, the description adds critical update-specific meaning not fully captured by the schema: omitted threshold/tag behavior and the need to send the full intended configuration. The schema's per-parameter details are already strong, and the tool description fills the omission/replacement gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with 'Replace an existing Monitor's exposed configuration' and enumerates the exact fields: name, Script, frequency, engine ids, failure and recovery thresholds, tags, and acceptable-metric thresholds. The word 'existing' distinguishes it from create_monitor, and the statement that it never changes whether the Monitor is enabled separates it from disable_monitor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly directs the caller to 'read the current configuration with get_monitor first and send everything that should remain' because omitted thresholds are cleared and omitted tags are removed. It also states the when-not: it cannot stop a running Monitor, start a stopped one, or alter notification policies or maintenance windows, so it should not be chosen for those intents.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_scenarioUpdate ScenarioADestructiveInspect
Replace an existing scenario's name and Populations — the populations given become the whole scenario, so to change one, send the others alongside it, each with the id get_scenario reported, or they are dropped and the ones you send are new populations. The same save rules and summary as create_scenario apply.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The new scenario name. | |
| projectId | Yes | ||
| scenarioId | Yes | ||
| populations | Yes | The Populations, in order. Each needs a name, the scriptId of a script in this project, and stages — { target, duration, ramp } segments that run for duration milliseconds and reach target bots by the end, each starting from where the previous stage left off. Optional per population: id (send back the id get_scenario reported to keep a population's identity), loadEngineId (an id from list_engines; defaults to the account's default region), iterations, iterationsPerUser, aggressionMultiplier, and playbackOptions with connectionBps (bits per second), hostnameOverrides, and variableOverrides. Read get_documentation topic 'load-test-scenarios' for what stages and ramps do; list_engines carries the account's stage bounds and bot caps. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description discloses exactly what gets destroyed: 'the populations given become the whole scenario... they are dropped and the ones you send are new populations.' This is a non-obvious, critical behavioral trait that prevents an agent from accidentally wiping out existing populations. References to shared save rules add further context, with no contradiction to annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loading the core action and then explaining the critical population-replacement semantics. It avoids restating schema details and uses the reference to create_scenario to avoid duplicating shared rules. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex mutation tool with no output schema, the description covers the essential invocation context: what it replaces, how to preserve populations, and where to find the ids. It relies on create_scenario for save rules and summary rather than specifying them, and does not mention projectId/scenarioId semantics, though those are inferable. This is highly functional but not fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, but the description compensates for the most complex parameter by explaining that populations are a full replacement and that existing identities must be sent back via the id from get_scenario. projectId and scenarioId have no schema descriptions, but their roles are self-evident from names; the description does not add detail for them, but the key behavioral parameter is well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific action 'Replace an existing scenario's name and Populations', identifying the resource (scenario) and what changes (name, populations). It clearly distinguishes from sibling create/delete/get operations by emphasizing replacement of an existing entity, and the whole-population caveat adds precision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: to change an existing scenario's name or replace its populations. It references get_scenario for ids and create_scenario for save rules, but it does not explicitly state when not to use it or contrast with create_scenario/delete_scenario as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_scriptUpdate ScriptADestructiveInspect
Replace an existing script's name, type, commands, and variables, creating a new revision. Every replace on this surface is a full replacement, not a merge, and every write against a user's existing data needs their explicit confirmation first.
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | The new script name, type, commands, and variables. | |
| scriptId | Yes | ||
| projectId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description explains that the whole script object is replaced rather than merged, that a new revision is created, and that writes to existing user data require explicit confirmation. These are meaningful behavioral disclosures an agent needs to invoke the tool safely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences: the first states the operation and result, the second gives the critical replacement and confirmation constraints. No filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool without an output schema, the description supplies the essential context: what is replaced, that a revision is created, that no merge occurs, and that user confirmation is required. The required identifiers and nested fields are already defined in the input schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 33% schema description coverage, the description compensates by enumerating the replaced fields (name, type, commands, variables) and framing them as a full replacement rather than a merge. It doesn't elaborate on projectId/scriptId, but their purpose is evident from names and the nested script schema covers the field structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific verb 'Replace' and the resource 'an existing script's name, type, commands, and variables', making the action and scope unmistakable. It also notes the result ('creating a new revision'), distinguishing it from create/delete operations without mentioning sibling names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for modifying an existing script via full replacement, but it does not explicitly state when to choose this tool over related alternatives such as create_script or duplicate_script, nor does it list conditions where it should not be used. The no-merge rule is behavioral context, not usage routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_scriptValidate ScriptARead-onlyIdempotentInspect
Structurally validate a list of commands for a bot type before saving or playing, reporting any issues by command index. Pass the script's variables as well and it also cross-checks the bindings against the ${...} and bot.getVariable reads in the commands, warning about a read nothing supplies, a binding nothing reads, and a per-occurrence variable read more than once in an iteration. Warnings never make a script invalid and never stop a save — all three are legal scripts.
| Name | Required | Description | Default |
|---|---|---|---|
| botType | Yes | The bot type the commands are for. | |
| commands | Yes | The commands to validate, each a JSON object with a 'type'. | |
| variables | No | The script's variable bindings, as get_script returns them. Omit to skip the binding cross-check and validate structure alone. |
Output Schema
| Name | Required | Description |
|---|---|---|
| valid | Yes | |
| issues | Yes | |
| warnings | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish readOnly, idempotent, and non-destructive behavior. The description adds meaningful behavioral detail beyond that: warnings never make a script invalid and never block a save, and it describes the specific warning categories. This helps an agent predict consequences accurately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the main purpose and is dense with useful information. The second sentence is long, but every clause conveys a distinct warning type or behavioral guarantee, so the length is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description does not need to explain return structure. It covers when to use the tool, what optional input changes behavior, and what warnings do and do not mean. No critical information for selecting and invoking the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents botType, commands, and variables well. The description reinforces the optional variables behavior and the cross-check logic, but it does not add substantially new parameter-level format or syntax information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'structurally validate a list of commands for a bot type', reporting issues by command index. It clearly distinguishes this from sibling tools like play_script or create_script by focusing on validation rather than execution or persistence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: validate 'before saving or playing' a script. It also explains when to include the optional variables parameter to enable the binding cross-check. It does not name alternatives like get_command_schema, but the intended timing and optional behavior are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
submit_feedback2 fields changed- changed
Input schema / properties / feedback / descriptionPrevious value: -"What worked and what did not, in at most 2000 characters."New value: +"What worked and what did not about the tools, in at most 2000 characters." - changed
Input schema / properties / taskDescription / descriptionPrevious value: -"What the task was, in at most 200 characters. Nothing on the server knows what you were doing, so this is the only context a reader has beyond the tool calls themselves."New value: +"The kind of task the note is about, in your own summary rather than the user's wording, in at most 200 characters."
55 tool updates
- First observed
append_dataset_rows - First observed
create_dataset - First observed
create_monitor - First observed
create_scenario - First observed
create_script - First observed
delete_dataset - First observed
delete_monitor - First observed
delete_scenario - First observed
delete_script - First observed
delete_script_asset - First observed
disable_monitor - First observed
duplicate_script - First observed
get_command_schema - First observed
get_dataset - First observed
get_documentation - First observed
get_example - First observed
get_incident - First observed
get_load_test_report - First observed
get_monitor - First observed
get_monitor_cycle_detail - First observed
get_monitoring_summary - First observed
get_play_status - First observed
get_scenario - First observed
get_screenshot - First observed
get_script - First observed
get_script_asset - First observed
get_script_revision - First observed
get_scripting_api - First observed
get_step_detail - First observed
get_variable_schema - First observed
import_script - First observed
list_command_types - First observed
list_datasets - First observed
list_engines - First observed
list_incidents - First observed
list_load_tests - First observed
list_monitor_cycles - First observed
list_monitoring_locations - First observed
list_monitors - First observed
list_projects - First observed
list_scenarios - First observed
list_script_assets - First observed
list_script_revisions - First observed
list_scripts - First observed
play_script - First observed
put_script_asset - First observed
restore_script_revision - First observed
stop_script - First observed
submit_feedback - First observed
update_dataset - First observed
update_load_test_notes - First observed
update_monitor - First observed
update_scenario - First observed
update_script - First observed
validate_script
Related MCP Connectors
Load & browser performance testing — drive MaxoPerf from your AI agent with your API key.
Website uptime monitoring: run checks from 300+ locations, manage monitors, alerts and incidents
Agents that test your web and mobile app like real users, on every pull request.
Synthetic checks, nightly regression replay and model-drift alerts for AI agents
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables automated browser testing of web applications using Playwright, supporting user interactions, form submissions, console monitoring, network request inspection, and visual verification through screenshots.-
- AlicenseBqualityCmaintenancePlaywright browser automation tuned for single-page apps (React/Vue/Angular) — screenshots, form filling, action chains, persistent sessions and 140+ device presets.22MIT
- AlicenseAqualityBmaintenanceEnables AI assistants to create, run, autocorrect, and monitor Playwright tests for your app with a local dashboard, credential vault, and public status page.16MIT

ProdPoke MCP Serverofficial
AlicenseNot gradedqualityDmaintenanceEnables AI to perform real-browser QA testing on websites via Playwright, finding bugs, accessibility, SEO, and performance issues through natural language conversations.1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.