jmeter-mcp-server
This server gives an LLM deterministic control over Apache JMeter: create, edit, run, and analyze JMeter load tests through typed MCP tools without writing JMX.
Author test plans: create TestPlan, thread groups (regular/setUp/tearDown), HTTP/JDBC/JSR223/FTP/TCP samplers, logic controllers, config elements, timers, pre/post-processors, assertions, and listeners.
Edit plans by id: remove, update, rename, move, reorder, enable/disable elements on a JSON tree; import existing
.jmx(unknown elements preserved) and export to JMX.Run real load tests asynchronously: execute a plan, poll status/progress/log, stop runs, and get aggregated per-label/overall stats (error %, avg/min/max/median/p90/p95/p99, throughput, KB/sec).
Find capacity limits: run automated breaking-point searches that ramp concurrency against SLA thresholds (p95/error %) and report the healthy/broken bound.
Inspect workspace: list plans, view full element trees, and serialize any plan to
.jmxXML.
Allows building, maintaining, running, and reading reports for Apache JMeter test plans without the GUI, including creating test plan elements, executing non-GUI load tests, and retrieving aggregated latency/error statistics.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jmeter-mcp-serverCreate a test plan hitting https://api.example.com/login with 50 users and show me the report."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jmeter-mcp-server
Give an LLM real, deterministic control over Apache JMeter — build test plans, run real load tests, and read back real results, without ever hand-writing .jmx XML or opening the GUI.
"Load test my API: 20 users hitting POST /orders for 2 minutes,
5% think time, fail anything over 800ms"...turns into a running JMeter test and a real report, through typed tool calls an MCP client (Claude Code, Claude Desktop, etc.) makes directly.
Why not just ask an LLM to write the .jmx itself?
It can — a .jmx is just XML, and any capable model has seen plenty of JMeter test plans. The problem is how it fails: JMeter's format is a hashTree with dozens of fragile, easy-to-misremember details — exact guiclass/testclass pairs, property names that don't match their GUI label (ThreadGroup.num_threads is a stringProp, not an intProp), integer bitmasks for assertion match types, strict parent/child pairing with sibling <hashTree> tags. None of it is self-checking. A wrong value still produces valid, loadable XML that just quietly does the wrong thing.
That's not hypothetical — it happened building this project. An early version of the If Controller generated a property called useExpression set to true, which reads like "yes, evaluate my condition." The real JMeter source does the opposite: useExpression=true means don't evaluate it as an expression — just check if the string is literally "true". Every non-trivial condition silently, permanently failed. No error, no warning — the child sampler just never ran. It only surfaced by actually executing the generated plan against real JMeter and noticing a sample count of zero.
That's the whole case for this server in one story: an LLM regenerating XML from memory re-risks that exact mistake on every single request. This server encodes the correct shape once, in a serializer (and a matching parser for the reverse direction) checked against real JMeter source and bundled examples, exercised by 166 automated tests including real JMeter runs — and exposes it as typed tools instead. Concretely:
Correctness through one tested code path, not regenerated-from-memory XML every time.
Cheap incremental edits. Plans are a small JSON tree with stable node ids — adding, removing, renaming, moving, or disabling an element is one tool call by
id, not rewriting a whole.jmxfile. An existing.jmx(hand-written or exported from the GUI) can be imported and edited the same way.Aggregated results, not raw samples.
get_execution_reportreturns computed stats (error %, avg/median/p90/p95/p99, throughput) — not thousands of sample rows to average by hand.Real async execution.
execute_test_planreturns immediately with anexecutionId; long-running load tests never block anything.
The generated .jmx is standard JMeter output — open it in the real GUI any time.
Related MCP server: JMeter MCP Server
Example
You: Build a load test: 10 users for 30s hitting GET https://api.example.com/health,
fail anything that takes over 500ms, then run it and tell me the p95.
Claude: [create_test_plan, add_thread_group, add_http_sampler, add_duration_assertion,
add_aggregate_report_listener, execute_test_plan, get_execution_status, get_execution_report]
Ran 300 requests over 30s, 0 failures. p95 latency: 214ms, avg: 187ms, throughput: 10.1 req/s.Every step above is a real typed MCP tool call — see Tools for the full set (35 element types across samplers, controllers, timers, extractors, assertions, and listeners, plus editing, inspection, and .jmx import/export tools) and Example workflow for the raw call sequence.
Quick start
claude mcp add jmeter \
-e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
-- npx -y jmeter-mcp-serverThat's it — no cloning, no build. Adjust JMETER_HOME to your JMeter install (see Prerequisites). Full setup details, Claude Desktop config, and local-dev instructions are in Adding this server to Claude Code.
How a test plan is represented
Each plan is a JSON tree ({id, type, props, children[]}), not XML text. Authoring tools append a child under a given parentId, editing tools (remove_element, update_element, move_element, etc.) mutate that same tree in place, and the tree is only serialized to a real .jmx file on demand (get_test_plan_xml) or at execution time. import_test_plan runs the reverse direction, parsing an existing .jmx back into this same tree shape. This is what makes incremental edits cheap and keeps all the fiddly XML schema knowledge in two places (src/jmx/serializer.ts for tree → XML, src/jmx/parser.ts for XML → tree, sharing prop shapes from src/jmx/propTypes.ts) instead of spread across every tool.
Tools
Authoring (each returns the new node's id, used as parentId for
whatever you attach under it next) — grouped the same way JMeter's own
right-click Add menu groups them, so if you already know the GUI, you
already know where to look:
Tool | Adds |
| Root |
Threads (Users):
Tool | Adds |
| Thread Group (virtual users); optional |
| setUp Thread Group (runs once before all Thread Groups) |
| tearDown Thread Group (runs once after all Thread Groups) |
Sampler:
Tool | Adds |
| HTTP Request sampler |
| JDBC Request sampler |
| JSR223 Sampler (Groovy/BeanShell/JS/JEXL script as the sample) |
| FTP Request sampler |
| TCP Sampler |
Logic Controller:
Tool | Adds |
| Transaction Controller (groups child samplers into one named transaction) |
| Loop Controller (repeats child samplers) |
| If Controller (conditionally runs child samplers) |
| While Controller (repeats children while a condition holds) |
| Random Controller (runs one random child per pass) |
| Interleave Controller (alternates through children) |
| Once Only Controller (children run only on each thread's first iteration) |
Config Element:
Tool | Adds |
| CSV Data Set Config (parameterization from a file) |
| User Defined Variables |
| JDBC Connection Configuration (pooled datasource) |
| HTTP Request Defaults |
| HTTP Cookie Manager |
| HTTP Header Manager |
Timer:
Tool | Adds |
| Constant Timer (pacing/think-time) |
| Uniform Random Timer (randomized pacing) |
| Constant Throughput Timer (target rate pacing) |
Pre Processors:
Tool | Adds |
| JSR223 PreProcessor |
| User Parameters (per-thread variable value sets) |
Post Processors:
Tool | Adds |
| JSON Extractor post-processor |
| Regular Expression Extractor post-processor |
| XPath Extractor post-processor |
| JSR223 PostProcessor |
Assertions:
Tool | Adds |
| Response Assertion |
| JSON Assertion (JSONPath validation) |
| Duration Assertion (response-time SLA) |
| Size Assertion (response byte-size check) |
Listener:
Tool | Adds |
| Aggregate Report listener |
| Summary Report listener |
| View Results Tree listener (full request/response capture for debugging) |
| Backend Listener (streams live metrics to InfluxDB/Graphite/etc.) |
Editing (mutate an already-built plan):
Tool | Purpose |
| Remove an element (and its subtree); rejects removing the root |
| Shallow-merge (or replace) a node's props; a prop value of |
| Rename an element's |
| Move an element (and its subtree) to a new parent, optionally at a specific index; rejects moving a node into its own subtree |
| Reorder a node's direct children (must pass an exact permutation of the current children) |
| Enable/disable an element without removing it |
Inspection:
Tool | Purpose |
| List every plan in the workspace |
| Full element tree of a plan, including every node's |
| Serialize a plan to its JMeter |
| Import an externally authored |
Execution & reporting (async — a run happens in the background):
Tool | Purpose |
| Serialize to |
|
|
| Send |
| Aggregated stats (per label + overall) parsed from the run's JTL output |
Capacity search (async — the search runs many real test executions in the background):
Tool | Purpose |
| Find the concurrency level where a plan stops meeting its SLA, by running it over and over and driving one thread group's thread count; returns |
|
|
| Abort a running search, keeping the bounds it had already established |
Finding the breaking point
find_breaking_point answers the question JMeter itself has no feature for: how
many concurrent users can this survive? Instead of you adjusting the thread
count, re-running, reading the report and deciding again — roughly six to eight
tool calls per round, across the six to eight rounds a search needs — the server
runs the whole search itself and you poll one tool for the answer.
It works in two phases:
Bracketing — starts at
startThreadsand doubles the load (50 → 100 → 200 …) while the SLA holds, until a round breaks it or themaxThreadsceiling is reached.Binary search — bisects between the last healthy level and the first broken one until the two are within
toleranceThreadsof each other, ormaxIterationsrounds have run, whichever comes first.
find_breaking_point (planId, threadGroupNodeId, maxThreads: 400, p95Ms: 800, errorPct: 1)
→ { searchId }
get_breaking_point_status (searchId) ← poll
→ { breakingPoint: 137, lastHealthy: 134,
breakingPointRange: { healthyUpTo: 134, brokenAt: 137, exact: false },
stopReason: "converged", progress: {...}, iterations: [...] }breakingPoint is the lowest load that was actually tested and broke the SLA.
Because the search stops bisecting at toleranceThreads, the levels between
lastHealthy and breakingPoint were never run — read breakingPointRange for
the real precision (exact: true only when the two are adjacent), and treat the
edge as "between 134 and 137", not "137". Status also reports progress (rounds
done, and the in-flight round's elapsed time and percent complete) and files
(where meta.json and the per-round executions live). There is no completion
push: poll get_breaking_point_status.
A round passes when every threshold set (p95Ms, errorPct) is met by the
overall TOTAL row; at least one threshold is required. Each round puts the
thread group into scheduler mode — a ramp-up proportional to the thread count
(rampSecondsPerThread) followed by a fixed plateauDurationSeconds at full
load — and only the plateau samples count, so the ramp-up doesn't drag the
numbers of a healthy round down. Loop counts are deliberately not used: they
would make rounds at different thread counts incomparable.
Because a round repeats the thread group's whole scenario until the plateau
ends, the request mix per round follows from how the plan is built: a login
that should happen once per user belongs under a Once Only Controller
(add_once_only_controller), or it will run on every repeat. Each round reports
its metrics overall and per label (byLabel, with each label's share of the
samples), so a skewed mix is visible right away — the SLA itself is judged on
the overall numbers.
The search temporarily overwrites numThreads, rampTimeSeconds,
durationSeconds and loops on the thread group you point it at, and restores
the original values when it ends — including when it fails or is stopped. The
status reports this as propsRestored, and the original values are kept in the
search's meta.json in case the process dies mid-search. The plan needs an
Aggregate Report, Summary Report, or View Results Tree listener, or there would
be no metrics to judge the SLA against.
Parameter | Default | Purpose |
| (required) | Safety ceiling — the search never runs more threads than this |
| (at least one required) | SLA thresholds a round is judged against |
|
| Load for the first round |
| 2% of | How close the bounds must get before the search stops bisecting |
|
| Ramp-up seconds per thread, so every round adds load at the same rate |
|
| Seconds at full load; only these samples count toward the SLA |
|
| Pause between rounds so the system under test recovers |
|
| Hard cap on rounds, so a slow search still ends |
A search that never breaks the SLA reports breakingPoint: null with
stopReason: "ceiling-reached" — raise maxThreads and go again. One where
even the first round breaks reports lastHealthy: null; lower startThreads.
While an execution runs, get_execution_status also returns a progress block
computed from the results file so far — samples, error rate, avg/p95 latency,
overall and recent throughput, and the same per label — so polling it is enough;
there is no need to tail stdout.log. Once the run ends, get_execution_report
gives the full aggregate.
JSR223 scripts need a compatible Java
JSR223 elements default to Groovy, and the Groovy bundled with JMeter 5.6.x
(3.0.x) fails on very new Java releases: every script run throws, which ends
that virtual user's iteration and can look like load that never grows. Use Java
17 (LTS) for JMeter — set JAVA_HOME in the MCP server's environment. The
add_jsr223_* tools, execute_test_plan and find_breaking_point detect the
Java and Groovy versions in use and return a warning / warnings entry when
they don't match, and the tool descriptions steer clients toward the built-in
${__UUID}, ${__RandomString} and ${__Random} functions, which need no
script.
Example workflow
create_test_plan → { planId, rootNodeId }
add_thread_group (parentId: rootNodeId) → { nodeId: threadGroupId }
add_http_sampler (parentId: threadGroupId) → { nodeId: samplerId }
add_response_assertion (parentId: samplerId)
add_aggregate_report_listener (parentId: threadGroupId)
execute_test_plan (planId) → { executionId }
get_execution_status (executionId) ← poll until "completed"
get_execution_report (executionId) → aggregated latency/error statsTesting
212 automated tests, no framework beyond Node's built-in test runner:
npm test # 196 tests: tree-mutation and XML-shape unit tests, XML -> tree parsing,
# serialize -> parse round-trips, the breaking-point search algorithm
# driven round-by-round against a simulated system, and every tool called
# over the real MCP protocol (stdio, the same way Claude Code/Desktop talk
# to it) - no JMeter install needed, fully hermetic
npm run test:integration # 16 tests: real JMeter runs - the If Controller story above, While
# Controller loop counts, timer pacing, extractors, assertions, and a full
# find_breaking_point search converging on the real concurrency limit of a
# live service (needs JMETER_HOME)
npm run test:allnpm test spawns the actual built server (dist/index.js) via StdioClientTransport and drives it exactly as a real client would — not just calling internal functions — so a broken tool schema or a malformed response shows up as a real protocol error, not a passing unit test.
Both suites run on every push and pull request via GitHub Actions — the integration job installs a real JMeter binary on the runner, so it's exercising the same code path as a local run, not a mock.
Prerequisites
Node.js 18+
JMeter installed locally, with the
JMETER_HOMEenvironment variable pointing at the installation directory (the one containingbin/jmeter). On macOS via Homebrew,brew install jmeterputs it at/opt/homebrew/opt/jmeter/libexec.
Adding this server to Claude Code
Via npx (recommended — published on npm)
No cloning or building required; npx fetches and runs the published
version on the fly:
claude mcp add jmeter \
-e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
-- npx -y jmeter-mcp-serverAdjust the JMETER_HOME path to wherever JMeter is installed on your
machine. Optionally set JMETER_MCP_WORKSPACE too (see below) if you want
plans and executions stored somewhere other than the default.
The default scope is local (this project directory only). To make it
available across every project, add -s user:
claude mcp add jmeter -s user \
-e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
-- npx -y jmeter-mcp-serverConfirm it registered and is responding:
claude mcp listFrom a local clone (development)
If you're working on this repository's code instead of using the published
package, point at the built dist/index.js directly:
npm install
npm run build
claude mcp add jmeter \
-e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
-- node /absolute/path/to/jmeter-mcp-server/dist/index.jsClaude Desktop
Add this to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"jmeter": {
"command": "npx",
"args": ["-y", "jmeter-mcp-server"],
"env": {
"JMETER_HOME": "/opt/homebrew/opt/jmeter/libexec"
}
}
}
}Note: unlike a terminal-launched app, Claude Desktop does not inherit
environment variables exported in your shell profile (.zshrc, etc.) — only
true system-wide ones. Always set JMETER_HOME explicitly in the env
block above rather than relying on it already being "set on your machine".
Environment variables
Variable | Required | Purpose |
| Yes | JMeter installation directory (must contain |
| No | Where plans and executions are stored. Defaults to |
Workspace layout
<workspace>/
plans/<planId>/plan.json # JSON tree — source of truth for a plan
executions/<executionId>/
generated.jmx # serialized at execute_test_plan time
aggregate-report.jtl # output of the Aggregate Report listener, if present
summary-report.jtl # output of the Summary Report listener, if present
jmeter.log
meta.json # execution status, pid, timestamps, exit code
capacity-searches/<searchId>/
meta.json # find_breaking_point search: config, SLA, every round run,
# current bounds, and the thread group's original settingsA find_breaking_point search doesn't get its own execution directory — each of
its rounds is a normal execution under executions/, and the search's
meta.json records the executionId of every round, so any individual round can
still be inspected with get_execution_report.
Editing and importing plans
Beyond the add_* authoring tools, a plan can be mutated after the fact
(remove_element, update_element, rename_element, move_element,
reorder_children, set_element_enabled) and an externally authored .jmx
(e.g. exported from the JMeter GUI) can be brought in with import_test_plan.
import_test_plan understands the most common element types (thread groups,
HTTP samplers, assertions, extractors, controllers, config elements, the
three report listeners, etc.); anything it doesn't recognize is kept as an
opaque UnknownElement node whose original XML is preserved and re-emitted
as-is by get_test_plan_xml/execute_test_plan, instead of being dropped -
import_test_plan's response reports unknownElementCount/
unknownElementTypes so you know what wasn't fully understood. Coverage can
be extended incrementally in src/jmx/parser.ts.
v1 scope
Not yet supported (candidates for a future release): generating the HTML
dashboard report (-e -o), parent-type validation on add_*/move_element/
import_test_plan (nothing stops attaching an element under a semantically
wrong parent), distributed execution.
Note on add_csv_data_set: the filename must be an absolute path.
execute_test_plan runs JMeter from a fresh per-execution directory, so a
relative path (which JMeter's GUI would resolve against the .jmx file's
own location) won't resolve there. An absolute path baked into a plan is
also machine-specific — it won't travel if you share plan.json with
someone on a different machine. This is only checked at creation time:
add_csv_data_set rejects a relative or nonexistent path up front, but
later changing a CSVDataSet's filename via update_element, or
importing a .jmx that already has a relative one via import_test_plan,
is not checked - it will only surface as a failure at execute_test_plan
time.
Note on add_jdbc_request/add_jdbc_connection_configuration,
add_ftp_request, and add_backend_listener: these generate correct,
JMeter-loadable XML, but exercising them for real needs infrastructure this
project doesn't provide (a database, an FTP server, an InfluxDB/Graphite
instance) — they were verified structurally, not against a real backend.
Note on add_view_results_tree_listener's captureFullData option: it has
no effect right now. execute_test_plan always runs JMeter with
-Jjmeter.save.saveservice.output_format=csv, and JMeter's CSV writer never
emits response body/header columns no matter what the SampleSaveConfiguration
flags say — only its XML output format can carry full response bodies. The
option is wired up correctly in the generated .jmx (verified: the flags
really do flip in the XML) for the day this server supports XML-format runs,
but until then it's a no-op — confirmed by running a real capture and
checking the resulting JTL has no responseData/samplerData/
requestHeaders/responseHeaders columns regardless of the setting.
Note on add_tcp_sampler: server/port/request are live-verified. The
numeric fields (connectTimeoutMs, timeoutMs) are rendered as
stringProp following this project's general convention for sampler
numeric fields, but that specific choice for TCPSampler wasn't confirmed
against a real JMeter-GUI-saved example (none was available to check
against) — flagging in case a real save turns out to expect intProp.
Roadmap
Ideas being explored for future releases — none of these are implemented yet:
Proposed tool | What it does | Why it's worth it |
| Fits Little's Law / the Universal Scalability Law to collected concurrency vs. throughput vs. latency data, and classifies the bottleneck as contention, coherency, or saturation | JMeter only outputs raw numbers; re-deriving a queueing-theory curve fit through prose reasoning would mean reimplementing nonlinear regression by hand for every question |
| Runs linear regression over the latency/error time series of a long-duration soak test to separate normal noise from a real trend (the classic memory-leak signal) | JMeter's graph shows the curve, but doesn't say whether it's a statistically real degradation or just noise |
| Statistical diff across N executions (not just two), with a significance test for whether a p95 shift is real or noise | A ready-made performance regression gate for CI, instead of someone eyeballing two JSON reports and guessing |
| Automatically detects where warm-up (JIT, connection pools, cold caches) ends, and recomputes metrics only over the steady-state window | Today's overall average is polluted by the first few seconds of a run; JMeter doesn't separate this on its own |
| Re-runs failed samples with backoff and separates "real system error" from "one-off flake" (network timeout, etc.) | Produces a trustworthy error rate for a CI gate, instead of an |
| Runs the same test plan across multiple JMeter injector nodes (master-slave, via | JMeter supports distributed mode, but wiring up remote engines, RMI ports, matching JMeter versions, and merging results across machines is entirely manual today — nobody sets this up for a quick test |
How this compares
Checked against the other public JMeter MCP servers found on GitHub as of September 3, 2026 — open-source repositories with at least a README description (undocumented forks/clones excluded). Columns reflect what each project's own README documents, not independent verification of its internals.
Server | Approach |
| Edit by id | Async run | Tests | Real JMeter |
jmeter-mcp-server (this project) | JSON tree, 34 element types, authored and edited by node id | ✓ | ✓ | ✓ | 212, incl. real runs | ✓ |
Runs an existing | — | ✗ | ✗ | ✗ | ✓ | |
Generates a | ✗ | ✗ | ✗ | ✗ | ✓ | |
Generates a whole plan from one JSON payload, runs it via Docker | ✗ | ✗ | ✗ | ✗ | ✓ | |
Heals the local JMeter runtime, imports HAR/OpenAPI traffic, discovers capacity | — | ✗ | ✗ | ✗ | ✓ | |
Demo project for a tutorial video, 11 fixed tools | ✗ | ✗ | ✗ | ✗ | ✓ | |
Reimplements HTTP load testing in Node — no Apache JMeter underneath | ✗ | ✗ | ✗ | ✗ | ✗ | |
Generates a | ✗ | ✗ | ✗ | ✗ | ✓ | |
Demo script with parameters hardcoded in source, not a general-purpose server | ✗ | ✗ | ✗ | ✗ | ✓ |
Every capability in that table shows up somewhere across the other eight projects, individually:
real JMeter execution, .jmx generation, traffic import. This is the only one that combines full
.jmx round-trip (import and export), per-id incremental editing, non-blocking async execution,
and a documented automated test suite in one place.
License
MIT
Available Tools
56 toolsadd_aggregate_report_listenerA
Add an Aggregate Report listener under the given parent (Thread Group or TestPlan). Its output is what get_execution_report reads after a run.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Aggregate Report | |
| planId | Yes | ||
| filename | No | Filename kept in the .jmx for portability; ignored at execution time | |
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry full behavioral transparency. It only indicates an 'add' operation without disclosing side effects (e.g., modifying the plan), requirements, reversibility, or safety considerations. The relationship to get_execution_report is mentioned, but no side effects are revealed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core action and the key dependency are stated directly and in a logical order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides the essential role (feeds get_execution_report) and parent context, but lacks detail on required parameters (planId, parentId remain ambiguous) and does not disclose mutation side effects. For a tool with 4 parameters and no annotations, it is not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, so the description must compensate. It clarifies parentId (Thread Group or TestPlan) and the tool's purpose, but does not explain planId or name beyond the schema's default, nor does it add much to filename. Compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (add) and a resource (Aggregate Report listener), and specifies the parent types (Thread Group or TestPlan). It also links the listener's output to get_execution_report, which distinguishes it from other listener tools like add_summary_report_listener.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (to produce output for get_execution_report), but it does not explicitly name alternatives or provide when-not-to-use guidance. The context is helpful but not as explicit as it could be.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_backend_listenerA
Add a Backend Listener under the given parent (Thread Group or TestPlan), streaming live metrics to an external backend (InfluxDB by default) instead of a JTL file. If args is omitted, a ready-to-edit InfluxDB argument set is used - at minimum, edit influxdbUrl before running.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | ||
| name | No | Backend Listener | |
| planId | Yes | ||
| parentId | Yes | ||
| classname | No | org.apache.jmeter.visualizers.backend.influxdb.InfluxdbBackendListenerClient |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the streaming behavior, the InfluxDB default, the default-args fallback, and even warns that influxdbUrl must be edited before running. It stops short of describing side effects or prerequisites beyond that, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the purpose front-loaded and the essential default-args caveat in the second. Every sentence earns its place and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the tool's core purpose and one important readiness caveat, but it leaves required planId semantics and configuration details under-specified. It is adequate, not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies parentId ('under the given parent') and args (omitted defaults to ready-to-edit InfluxDB arguments), but the required planId is not explained, and name/classname are left to their schema defaults without semantic context. Major parameter semantics are missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete action ('Add a Backend Listener'), the target location ('under the given parent (Thread Group or TestPlan)'), and the distinguishing behavior ('streaming live metrics to an external backend ... instead of a JTL file'). This differentiates it from sibling listener-adding tools and makes its role immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when this tool is appropriate: when live metrics should go to an external backend such as InfluxDB rather than into a JTL file. It does not explicitly name alternative tools or state when-not-to-use, but the 'instead of a JTL file' contrast gives enough steering.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_constant_throughput_timerA
Add a Constant Throughput Timer under the given parent, pacing requests to hit a target rate rather than a fixed per-sample delay.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Constant Throughput Timer | |
| planId | Yes | ||
| calcMode | No | this_thread_only | |
| parentId | Yes | ||
| targetSamplesPerMinute | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It explains the core behavior (pacing to a target rate) and the placement under a parent, but it omits any side effects, prerequisites, or implications of calculator mode. This is adequate for a basic add operation, but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly-worded sentence with no filler. It front-loads the action and resource, then adds the differentiating behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With five parameters, zero schema description coverage, no annotations, and no output schema, the description is too thin to be fully complete. An agent would need to consult the schema for required fields and enum semantics, and the description offers little contextual help about when or how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not name or explain any of the parameters. It vaguely references 'target rate' but does not map that to targetSamplesPerMinute, nor does it explain planId, parentId, name, or calcMode. The burden falls entirely on the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Add'), a specific resource ('Constant Throughput Timer'), and a specific placement ('under the given parent'). It also distinguishes itself by contrasting target-rate pacing with a 'fixed per-sample delay', which clearly separates it from sibling timers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this tool when you need throughput-based pacing rather than a fixed per-sample delay. It does not explicitly name alternatives like add_constant_timer or add_uniform_random_timer, but the 'rather than' phrasing implies the comparison, so it stops just short of explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_constant_timerA
Add a Constant Timer under the given parent to pace requests. Under a Thread Group it delays every sampler in it; under a single sampler it delays only that one, before it fires.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Constant Timer | |
| planId | Yes | ||
| delayMs | Yes | Milliseconds to pause before each sample | |
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the burden of behavioral disclosure. It does so well by explaining that under a Thread Group it delays every sampler and under a sampler it delays only that sampler 'before it fires.' It does not disclose side effects like whether an existing timer is replaced or what the response contains, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, tightly written, with no filler. The main action is front-loaded, and the scoping nuance follows in the second sentence. Every clause adds information needed to invoke the tool correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a relatively simple creation tool, the description is complete enough: it names the parent placement effect, the delay unit, and the timing of the delay. It does not explain planId explicitly or mention the response, but with no output schema and a simple mutation, the essential calling context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, with delayMs already documented in the schema. The description adds useful meaning for parentId by explaining how parent type changes scope, but it gives no additional guidance for planId or name. This partially compensates for low schema coverage but does not fully close the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Add a Constant Timer under the given parent to pace requests.' It then clarifies the exact behavioral scope for Thread Group vs. single sampler parents, which distinguishes it from sibling timer tools like add_uniform_random_timer and add_constant_throughput_timer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for use: it tells the agent what happens when placed under a Thread Group versus under a single sampler, and that its purpose is pacing requests. It does not explicitly enumerate alternatives such as 'use add_uniform_random_timer for randomized delays,' so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_cookie_managerA
Add an HTTP Cookie Manager under the given parent (usually a Thread Group), so samplers in scope share cookies automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | HTTP Cookie Manager | |
| planId | Yes | ||
| policy | No | standard | |
| parentId | Yes | ||
| clearEachIteration | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the main behavioral trait: adding the manager makes samplers in scope share cookies automatically. However, it does not mention how policy or clearEachIteration affect behavior, whether existing cookies are cleared, or any reversibility concerns. The core side effect is clear, but deeper behavioral context is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the action and resource, and no wasted words. Every clause adds either placement or behavioral value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no annotations, no output schema, and zero parameter documentation, the description is too thin. It covers the parent relationship and cookie-sharing effect, but fails to describe required planId, the policy options, clearEachIteration behavior, or what a successful call returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only implicitly explains parentId via 'given parent'; planId, name, policy, and clearEachIteration receive no explanation in either the schema or the description. This is a significant gap, though the parent reference earns some credit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add'), a specific resource ('HTTP Cookie Manager'), and the placement context ('under the given parent'). It also gives the operational outcome ('samplers in scope share cookies automatically'), which clearly distinguishes it from the many sibling add_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete placement guidance ('usually a Thread Group') and explains the functional consequence that makes the tool useful. It does not explicitly contrast with alternatives like add_header_manager or add_http_request_defaults, but the context is clear enough for an agent to decide when cookie management is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_csv_data_setA
Add a CSV Data Set Config under the given parent (a Thread Group to feed every sampler in it, or a single sampler to scope it there). Each thread reads one row per iteration, exposing each column as a JMeter variable. The filename must be an absolute path readable on the machine that will run the test (execute_test_plan runs JMeter from a fresh per-execution directory, so relative paths won't resolve).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | CSV Data Set Config | |
| planId | Yes | ||
| recycle | No | Start over from the first line once the last row is used | |
| filename | Yes | Absolute path to the CSV file | |
| parentId | Yes | ||
| delimiter | No | , | |
| stopThread | No | Stop the thread once the file is exhausted (instead of recycling) | |
| variableNames | No | Comma-separated column names, e.g. 'username,password'. Omit if the CSV has a header row. | |
| ignoreFirstLine | No | Set true if the CSV's first line is a header row |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It proactively discloses the critical requirement that filename must be an absolute path and warns about the per-execution directory logic in execute_test_plan, which is a non-obvious caveat. It also explains the row-reading behavior, which goes beyond the schema. This is strong, though it stops short of describing failure modes or the response format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place. The first states the core action and placement, the second conveys the primary behavioral mechanism, and the third delivers a crucial operational warning. No filler, no repetition of schema defaults. The most important caveat is saved for the end, but it's not overlong.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the description covers the essential placement, data-flow behavior, and the file-path pitfall. It does not explicitly address interactions between recycle and stopThread beyond what the schema already provides, nor does it mention whether the tool overwrites existing configs. Given the moderate complexity, the description is fairly complete but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 56%, so the description needs to compensate for undocumented parameters (name, planId, parentId, delimiter). It does by explaining the meaning of parentId (placement) and filename (absolute path required). It also adds context on how variableNames and ignoreFirstLine relate to CSV header handling implicitly. The schema already covers the other parameters, so the description adds value without duplicating schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Add a CSV Data Set Config') and names the parent as either a Thread Group or a single sampler, which clearly distinguishes it from sibling add_* tools like add_thread_group or add_http_sampler. It also explains the core behavior (each thread reads one row per iteration, columns exposed as JMeter variables), making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear placement guidance: attach under a Thread Group to feed all samplers, or under a single sampler to scope locally. This helps the agent decide where to attach. However, it does not explicitly state when to choose this over alternatives like add_user_parameters or add_user_defined_variables, nor does it mention scenarios where CSV config is inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_duration_assertionA
Add a Duration Assertion under an HTTP sampler, failing the sample if it takes longer than maxDurationMs.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Duration Assertion | |
| planId | Yes | ||
| parentId | Yes | ||
| maxDurationMs | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose the core effect ('failing the sample if it takes longer than maxDurationMs') and the placement. However, it does not state that this mutates the test plan, that the sampler must already exist, or what happens if parentId is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler. The action, placement, and behavior are front-loaded, and every phrase adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple add-operation with no annotations and no output schema, the description covers the main action and effect, but omits prerequisites (existing plan/sampler), mutation side effects, and error conditions. It is minimally adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains maxDurationMs (the duration threshold) and parentId (must be an HTTP sampler). It does not explain planId or the optional name parameter, leaving a gap for those parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Add'), names the resource ('Duration Assertion'), specifies placement ('under an HTTP sampler'), and states the behavioral effect ('failing the sample if it takes longer than maxDurationMs'). This clearly distinguishes it from sibling assertion tools like add_response_assertion and add_size_assertion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'under an HTTP sampler' gives clear placement context, telling the agent that parentId must reference an HTTP sampler. It does not explicitly name alternatives or exclusions, but the context is sufficient for this simple add-operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_ftp_requestC
Add an FTP Request sampler under the given parent (usually a Thread Group).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | FTP Request | |
| port | No | ||
| planId | Yes | ||
| server | Yes | ||
| upload | No | ||
| filename | Yes | Remote file path to download/upload | |
| parentId | Yes | ||
| password | No | anonymous@test.com | |
| username | No | anonymous | |
| inputData | No | Literal data to upload, instead of a local file | |
| binaryMode | No | ||
| saveResponse | No | ||
| localFilename | No | Local file path, for download destination or upload source |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden of behavioral disclosure. It only states the action and parent relationship, with no mention of upload versus download behavior, anonymous login defaults, side effects, or validation. An agent gets almost no behavioral guarantees beyond the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It earns its place by stating the core action and placement, though the brevity veers into under-specification rather than efficient completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 13 parameters, no output schema, and no annotations, one sentence is far from sufficient. An agent cannot safely construct a valid FTP request or understand the relationship between remote filename, local file path, and upload flag from this description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 23% across 13 parameters, and the description adds no parameter-level meaning. It does not explain server, filename, upload, inputData, localFilename, binaryMode, or any other field, so the agent must guess at how these parameters interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Add an FTP Request sampler under the given parent.' It clearly distinguishes this tool from the many sibling add_* tools by naming the exact sampler type and its intended placement context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies placement under a parent like a Thread Group, but it gives no guidance on when to choose this sampler over alternatives such as add_http_sampler, add_tcp_sampler, or add_jdbc_request. No prerequisites, exclusions, or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_header_managerB
Add an HTTP Header Manager under an HTTP sampler (or a Thread Group, to apply to every sampler in it).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | HTTP Header Manager | |
| planId | Yes | ||
| headers | Yes | ||
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states that it adds a header manager, but does not mention any side effects (e.g., modifying the test plan), required permissions, idempotency, or what happens if the parent doesn't exist. For a mutation tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no filler. It is front-loaded with the action and placement, making it easy to scan. However, the brevity contributes to the lack of parameter and behavior detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four parameters with zero schema descriptions, no output schema, and no annotations, the description should provide substantial guidance. It only covers the core purpose and placement, leaving out parameter semantics, required IDs, header structure, and any behavioral caveats. The tool is underspecified for reliable agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the four parameters (planId, parentId, headers, name). It fails to compensate for the schema's lack of documentation, leaving agents to guess the meaning and structure of the headers array and required IDs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (add an HTTP Header Manager) and specifies the target (under an HTTP sampler or Thread Group), which distinguishes it from sibling tools that add other elements. The placement detail adds precision beyond a generic 'add' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides placement guidance (under sampler vs. Thread Group) which implicitly tells when to use it for scoping headers to all samplers. It does not explicitly mention alternatives or when not to use it, but there are no similar sibling tools that add headers, so this is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_http_request_defaultsA
Add HTTP Request Defaults under the given parent (usually a Thread Group or TestPlan root). Any field left blank on a later add_http_sampler call under the same scope falls back to these values.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | HTTP Request Defaults | |
| path | No | ||
| port | No | ||
| domain | No | ||
| planId | Yes | ||
| parentId | Yes | ||
| protocol | No | ||
| connectTimeoutMs | No | ||
| responseTimeoutMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavior. It does: fields left blank on later add_http_sampler calls fall back to these values, and scoping is specified ('under the same scope'). However, it does not mention effects on existing samplers or behavior when defaults are empty, so it is not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The purpose is front-loaded, and the fallback behavior is explained in the second sentence. Excellent economy of language.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the core purpose and behavior, but given 9 parameters and no annotations, it lacks parameter-level detail (especially planId/parentId semantics) and does not mention prerequisites or impact on existing elements. The description is adequate but not fully complete for a tool with this many parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It only explains the collective fallback concept, not individual parameters like planId, parentId, or units for timeouts. Parameter names are somewhat self-explanatory but key identifiers and units are left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Add HTTP Request Defaults'. Clearly explains the fallback behavior that distinguishes it from add_http_sampler, and mentions typical placement (Thread Group or TestPlan root), making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly ties usage to add_http_sampler: defaults are used for later calls under the same scope. This gives clear context for when to use it, but does not explicitly state when not to use it or list alternative tools, so it's slightly below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_http_samplerA
Add an HTTP Request sampler under the given parent node (usually a Thread Group). To inherit protocol, domain or port from an HTTP Request Defaults config element in scope (add_http_request_defaults), leave that argument out of the call - do not pass an empty string or a placeholder, because those fields have no default of their own. Passing an explicit value always overrides the Defaults for that field.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| path | Yes | Request path, e.g. /v1/users | |
| port | No | Omit to inherit from HTTP Request Defaults | |
| domain | No | Host name, e.g. api.example.com. Omit to inherit from HTTP Request Defaults | |
| method | Yes | ||
| planId | Yes | ||
| bodyJson | No | Raw JSON request body, if any | |
| parentId | Yes | ||
| protocol | No | Omit to inherit from HTTP Request Defaults |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the inheritance behavior and the critical warning about not passing empty strings/placeholders, which is a behavioral trait beyond what the schema shows. However, it doesn't mention whether the operation is destructive or requires specific permissions, but for an add operation this is less critical. The description adds meaningful behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: it states the core action first, then the inheritance rule, then the warning. Every sentence earns its place, and the warning is critical for correct usage. No fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 params, 5 required, no output schema), the description covers the most important behavioral nuance (inheritance) and the key pitfall (empty strings). It doesn't explain return values, but with no output schema that's less critical. The description is complete enough for an agent to call the tool correctly, though it could mention that the sampler is added to a JMeter test plan context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 56%, so the description partially compensates. The description adds crucial semantics for protocol, domain, and port by explaining the inheritance behavior and the warning about empty strings. It also clarifies the parentId parameter ('usually a Thread Group'). However, it doesn't explain planId or bodyJson beyond what the schema already says, but the schema covers those adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: adding an HTTP Request sampler under a parent node, typically a Thread Group. It distinguishes itself from related tools by explicitly referencing add_http_request_defaults and explaining the inheritance relationship, which differentiates it from other add_* samplers like add_jsr223_sampler or add_tcp_sampler.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool and how to handle inheritance: omit protocol/domain/port arguments to inherit from HTTP Request Defaults, and avoid passing empty strings or placeholders. It also names the related tool add_http_request_defaults, giving clear context for when to use this vs. that tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_if_controllerA
Add an If Controller under the given parent (usually a Thread Group). Samplers added under it (as its children) only run when the condition evaluates truthy. condition is evaluated as a real JavaScript (Rhino) expression after JMeter substitutes any ${var} references, so comparisons work directly, e.g. '${count} < 10' or '"${status}" == "ok"'.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | If Controller | |
| planId | Yes | ||
| parentId | Yes | ||
| condition | Yes | JavaScript boolean expression evaluated after ${var} substitution, e.g. '${count} < 10' | |
| evaluateAll | No | Evaluate the condition for every iteration, not just the first |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that samplers run only when the condition is truthy and that the condition is evaluated as a JavaScript expression after variable substitution. However, it omits the default evaluation timing (first iteration only unless evaluateAll is true) and any side effects or failure behavior, which are important for a complete behavioral picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loaded with the primary action and purpose, and includes practical examples. Every sentence adds value without redundancy, making it efficient for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter creation tool with no output schema, the description covers the core purpose and condition semantics but misses key operational details such as the default evaluation behavior (evaluateAll=false), what the tool returns (e.g., created node ID), and any prerequisites or validations for parentId. This leaves the agent without full guidance for a correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, so the description must compensate for undocumented parameters. It adds substantial meaning for 'condition' with examples and clarifies the JavaScript/Rhino evaluation, but does not elaborate on 'planId', 'parentId', or 'name' beyond what the schema shows. It partially fills the gap but leaves other parameters to be inferred from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add') and resource ('If Controller'), clarifies placement under a parent, and explains the conditional execution of child samplers. It clearly distinguishes this tool from sibling controllers like add_while_controller or add_loop_controller by specifying the 'If' type and its condition-based behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for conditional execution but does not explicitly compare with alternatives or state when not to use it. It gives context (parent is usually a Thread Group) and condition evaluation details, but no direct guidance on choosing this over while/loop/random controllers. This leaves the agent to infer the appropriate use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_interleave_controllerA
Add an Interleave Controller under the given parent (usually a Thread Group). Runs one child (added under it) per pass, alternating through them in order instead of running them all.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Interleave Controller | |
| planId | Yes | ||
| parentId | Yes | ||
| ignoreSubControllerBlocks | No | If true, interleaves across every leaf sampler recursively instead of one-per-direct-child |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose the key behavioral trait – one child per pass, alternating in order – which is genuinely useful. However, it omits side effects, return value, whether children must already exist, and any permission or plan-modification details, so transparency is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and placement, followed by a compact behavioral explanation. No filler or repeated schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple add operation: it states what is added, where, and how it behaves, while the schema covers defaults and ignoreSubControllerBlocks. But with no output schema and no annotations, it leaves return value and explicit parameter semantics unstated, and it does not route to alternatives. Clear gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%; only ignoreSubControllerBlocks has a schema description. The description gives context for parentId ('under the given parent') but does not explain planId, name, or ignoreSubControllerBlocks, nor how they map to the interleave behavior. It therefore does not compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Add an Interleave Controller') and a placement target, then explains the distinguishing runtime behavior (one child per pass, alternating in order). This clearly separates it from sibling controllers like add_loop_controller or add_random_controller without needing to inspect their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives placement guidance ('usually a Thread Group') and implies the use case through the alternating behavior, but it never explicitly states when to choose this controller over siblings or when not to use it. No alternatives are named, so the agent must infer selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_jdbc_connection_configurationA
Add a JDBC Connection Configuration (pooled datasource) under the given parent, usually the Thread Group or TestPlan root. A JDBC Request references this element's dataSource name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | JDBC Connection Configuration | |
| dbUrl | Yes | JDBC URL, e.g. jdbc:postgresql://host:5432/dbname | |
| driver | Yes | JDBC driver class, e.g. org.postgresql.Driver | |
| planId | Yes | ||
| poolMax | No | Max pool size, 0 = unlimited | |
| timeout | No | ||
| parentId | Yes | ||
| password | No | ||
| username | No | ||
| checkQuery | No | ||
| dataSource | Yes | Pool name that JDBC Request samplers will reference | |
| trimInterval | No | ||
| connectionAge | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It explains the parent-child relationship and the dataSource linkage, which is useful behavioral context, but it does not disclose potential overwrites, validation behavior, required parent existence, or whether a connection is tested at creation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core action and placement are front-loaded, and the dataSource relationship is stated efficiently in the second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a 13-parameter creation tool with no annotations and no output schema, yet the description covers only the basic purpose and one relationship. Missing behavioral details, parameter guidance, and result expectations leave the agent under-informed for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low (31%), so the description should compensate for undocumented parameters. It only highlights the dataSource concept and parent placement; it adds no meaning for poolMax, timeout, checkQuery, trimInterval, connectionAge, username, password, or other parameters the schema leaves unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add'), a precise resource ('JDBC Connection Configuration (pooled datasource)'), and placement context ('under the given parent'). It also distinguishes the element from the sibling add_jdbc_request by explaining that a JDBC Request references this element's dataSource name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on where the element belongs ('usually the Thread Group or TestPlan root') and implies sequencing relative to JDBC Request samplers. It does not explicitly list exclusions or when-not-to-use conditions, but the intended usage is reasonably inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_jdbc_requestA
Add a JDBC Request sampler under the given parent (usually a Thread Group). dataSource must match a JDBC Connection Configuration's dataSource name added elsewhere in the same plan.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | JDBC Request | |
| query | Yes | ||
| planId | Yes | ||
| parentId | Yes | ||
| queryType | No | Select Statement | |
| dataSource | Yes | ||
| variableNames | No | Comma-separated variable names to store each result column under | |
| resultVariable | No | Variable name to store the whole result set under |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden. It discloses a key dependency (dataSource must match an existing configuration) and clarifies that the element is added to a parent. However, it does not cover failure behavior, return value, or what happens if the parent or dataSource is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The primary action is front-loaded, and the critical prerequisite is stated in the second sentence without unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no annotations and no output schema, the description covers placement and the main prerequisite but omits guidance on queryType semantics, required plan/parent state, and return behavior. It is adequate for selecting the tool but leaves several call-time details to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, so the description must compensate. It adds real meaning to dataSource (must match a JDBC Connection Configuration's name) and parentId (usually a Thread Group), but leaves query, queryType, and the variable storage parameters to their names and enum values. Compensation is partial but helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact element to create ('JDBC Request sampler'), the operation ('Add'), and the placement ('under the given parent'). It also distinguishes the tool from the sibling add_jdbc_connection_configuration by focusing on the sampler rather than the configuration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear prerequisite context: the sampler requires a JDBC Connection Configuration with a matching dataSource name in the same plan. It does not explicitly list alternatives or when-not-to-use conditions, but the parent placement and dependency are enough for an agent to select it appropriately among sampler siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_json_assertionB
Add a JSON Assertion under an HTTP sampler, to validate a JSONPath expression exists (and optionally matches a value) in the response.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | JSON Assertion | |
| invert | No | Negate the assertion | |
| planId | Yes | ||
| isRegex | No | Treat expectedValue as a regular expression | |
| jsonPath | Yes | JSONPath expression, e.g. $.value | |
| parentId | Yes | ||
| expectNull | No | ||
| expectedValue | No | Value to compare against, only checked if jsonValidation is true | |
| jsonValidation | No | Whether to check expectedValue at all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. It explains what the assertion validates but does not disclose that this mutates the test plan, any permission/side-effect implications, or behavior around failing validations/defaults. The optional-match behavior is mentioned but is closer to parameter semantics than to tool behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. It states the action, the target element, the placement condition, and the core validation behavior without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with nine parameters, no annotations, and no output schema, this description is too thin. It does not explain the required planId/parentId semantics, the meaning of expectNull, or what happens after a successful/failed addition; the agent would need to rely on incomplete schema text for several parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 56%, and the description adds a high-level hint that matching a value is optional, which aligns with jsonValidation/expectedValue. However, it does not clarify parameters like expectNull, planId, or parentId, and much of the parameter meaning still depends on the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Add a JSON Assertion') and states the tool's function: validating that a JSONPath expression exists and optionally matches a value. It is clear and points to the JSON assertion element, but it does not explicitly contrast it with sibling tools such as add_response_assertion or add_json_extractor, so sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'under an HTTP sampler' gives a placement constraint and the validation purpose implies when the tool should be used. However, there is no explicit guidance about when not to use it or how it differs from alternative assertion/extractor tools, leaving the agent to infer the choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_json_extractorA
Add a JSON Extractor post-processor under an HTTP sampler, to save a value from the JSON response into a variable.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| planId | Yes | ||
| parentId | Yes | Id of the HTTP sampler node this extractor applies to | |
| defaultValue | No | Value to use if the JSONPath doesn't match | NOT_FOUND |
| jsonPathExpr | Yes | JSONPath expression, e.g. $.data.id | |
| referenceName | Yes | JMeter variable name to store the extracted value in |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must fully disclose behavioral traits. It only says 'Add', implying a mutation, but does not explain side effects like whether existing extractors are replaced, what happens if the parent sampler does not exist, or whether the operation is reversible. This is a significant gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence that front-loads the action and purpose. Every word earns its place, with no redundant information or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is adequate but not complete. It explains the main use case, but does not mention potential behavioral details (e.g., failure handling, idempotency) or prerequisites beyond the parent sampler. The missing parameter descriptions for name and planId are not addressed. A more comprehensive description would help for a mutation operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% (4 of 6 parameters have descriptions). The description adds no additional parameter-specific details beyond what the schema already provides; it merely restates the purpose which maps to jsonPathExpr and referenceName. Since the schema covers most parameters, the baseline of 3 is appropriate, and the description does not compensate for the undocumented name and planId parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add'), a clear resource ('JSON Extractor post-processor'), a location context ('under an HTTP sampler'), and the purpose ('save a value from the JSON response into a variable'). This clearly distinguishes it from sibling add_* tools like add_http_sampler or add_response_assertion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use this tool: when you need to add a JSON extraction post-processor. It provides context about the parent sampler. However, it does not explicitly state when not to use it or mention alternatives (e.g., if you need a different extractor type), but given the sibling list, the purpose is distinct enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_jsr223_postprocessorA
Add a JSR223 PostProcessor under an HTTP sampler (or a Thread Group, to apply to every sampler in it), running a script after the sample. Scripts default to Groovy, which only works when JMeter runs on a compatible Java (17 is the safe choice; the Groovy bundled with JMeter 5.6.x fails on very new Java such as 25+, killing each virtual user on its first script run) - the result carries a "warning" field when this server detects a mismatch. For unique or random data, prefer the built-in functions ${__UUID}, ${__RandomString} and ${__Random}, which need no script at all.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | JSR223 PostProcessor | |
| planId | Yes | ||
| script | Yes | ||
| parentId | Yes | ||
| parameters | No | ||
| scriptLanguage | No | groovy |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that scripts default to Groovy, that Groovy can fail on very new Java (killing virtual users), and that the result carries a warning field on mismatch. This is rich behavioral disclosure beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and placement, then adds critical compatibility notes and an alternative. Each sentence earns its place; it is long but information-dense without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex scripting tool with no output schema, the description covers placement, language, failure modes, and an alternative. It omits the meaning of the 'parameters' string and doesn't state how the script's return value affects the sample, but core invocation details are present. Slightly incomplete on parameter semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions the scriptLanguage default (groovy) and that the script is the core, but does not explain planId, parentId, parameters, or the other language options. The description adds minimal parameter-level meaning beyond the schema defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool adds a JSR223 PostProcessor, specifies placement (under an HTTP sampler or Thread Group), and its function (running a script after the sample). It distinguishes itself from the sibling add_jsr223_preprocessor by explicitly noting 'after the sample'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit usage context: placement options, default script language, Java compatibility warning, and a direct alternative (built-in functions like ${__UUID}) for random data. It also implies when not to use a script, making selection guidance strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_jsr223_preprocessorA
Add a JSR223 PreProcessor under an HTTP sampler (or a Thread Group, to apply to every sampler in it), running a script before the sample. Scripts default to Groovy, which only works when JMeter runs on a compatible Java (17 is the safe choice; the Groovy bundled with JMeter 5.6.x fails on very new Java such as 25+, killing each virtual user on its first script run) - the result carries a "warning" field when this server detects a mismatch. For unique or random data, prefer the built-in functions ${__UUID}, ${__RandomString} and ${__Random}, which need no script at all.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | JSR223 PreProcessor | |
| planId | Yes | ||
| script | Yes | ||
| parentId | Yes | ||
| parameters | No | ||
| scriptLanguage | No | groovy |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses the pre-sample execution timing, the Groovy/Java compatibility risk that can kill virtual users, and the presence of a 'warning' field in the result. This goes beyond basic mutation and adds critical operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core purpose and usage. It includes a practical recommendation about built-in functions without padding. Slightly long but efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters and no output schema, the description covers the essential behavior, placement, and known pitfalls. It mentions the warning field and compatibility issues, which are important for correct invocation. Missing parameter explanations are a minor gap given the tool's context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies 'script' and 'scriptLanguage' (default groovy, with a compatibility caveat) but leaves 'planId', 'parentId', 'name', and 'parameters' unexplained. This is a partial improvement over the schema, but not sufficient for full parameter clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool adds a JSR223 PreProcessor, specifies where it can be attached (HTTP sampler or Thread Group), and describes its function (running a script before the sample). It also distinguishes it from related tools like postprocessor and sampler by the explicit 'PreProcessor' and timing context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use built-in functions instead of scripts, and warns about Java/Groovy compatibility. However, it does not directly compare against sibling tools like add_jsr223_postprocessor or add_jsr223_sampler, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_jsr223_samplerA
Add a JSR223 Sampler under the given parent (usually a Thread Group), running a script as the sample itself. Scripts default to Groovy, which only works when JMeter runs on a compatible Java (17 is the safe choice; the Groovy bundled with JMeter 5.6.x fails on very new Java such as 25+, killing each virtual user on its first script run) - the result carries a "warning" field when this server detects a mismatch. For unique or random data, prefer the built-in functions ${__UUID}, ${__RandomString} and ${__Random}, which need no script at all.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | JSR223 Sampler | |
| planId | Yes | ||
| script | Yes | ||
| parentId | Yes | ||
| parameters | No | ||
| scriptLanguage | No | groovy |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses a critical failure mode: Groovy's incompatibility with very new Java versions (25+), which kills virtual users, and mentions the resulting 'warning' field. This is valuable transparency. However, it doesn't mention other behaviors like whether the script is compiled, side effects on the test plan, or response format beyond the warning.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded. The first sentence states the core purpose, followed by a crucial compatibility warning, then a practical alternative. Each sentence earns its place, and the length is appropriate for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main purpose, the Java compatibility pitfall, and an alternative for common use cases. It mentions the 'warning' field in the result, which is helpful. However, it does not explain the 'parameters' field or the exact return value (no output schema), and it doesn't elaborate on prerequisites beyond the parent. Overall, it is sufficiently complete for an agent to call correctly in most scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'parent' (the parent element), 'script' (the script to run), and implicitly 'scriptLanguage' by mentioning the Groovy default. However, it does not explain 'parameters', 'name', or 'planId' at all. These are not described in the schema either, so the description only partially covers parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool adds a JSR223 Sampler under a parent and runs a script as the sample itself. It distinguishes this from sibling pre/post processors by specifying it is the sampler itself. The verb 'Add' plus the specific element name leaves no ambiguity about the action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance to prefer built-in functions like ${__UUID}, ${__RandomString}, and ${__Random} for unique or random data, thereby telling when NOT to use this tool. It also implies when to use it: when a scripted sampler is needed. This is a clear alternative and usage condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_loop_controllerA
Add a Loop Controller under the given parent (usually a Thread Group). Samplers added under it (as its children) repeat for the given number of loops each time the controller is reached.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Loop Controller | |
| loops | No | Number of iterations, or -1 to loop forever | |
| planId | Yes | ||
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the disclosure burden. It does reveal the key non-obvious behavior: samplers under the controller repeat for the specified loop count each time the controller is reached. However, it does not state whether the plan is modified immediately, how invalid parents are handled, or what the return value is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler: it leads with the action and target element, then explains the child-sampler repeat behavior. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple element-adding tool, the description covers placement and behavior, and the schema provides defaults and the -1 infinite-loop semantics. It is less complete on the required planId parameter and on return/error behavior, and with no output schema or annotations those gaps are not filled elsewhere. Overall it is adequate but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning for parentId ('under the given parent') and loops ('repeat for the given number of loops'), but planId—a required parameter—is never described or contextualized. With schema description coverage at only 25%, the description only partially compensates for missing parameter documentation; name is also left to its default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Add a Loop Controller'), the placement ('under the given parent'), and the resulting repeat behavior of child samplers. It is immediately recognizable among the many add_*_controller siblings by resource type and loop semantics. It does not explicitly contrast it with sibling controllers, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful placement context ('under the given parent, usually a Thread Group') and behavioral context (child samplers repeat by the loop count each time the controller is reached). This helps an agent decide when to pick this controller tool. It does not mention when not to use it or name an alternative controller.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_once_only_controllerA
Add a Once Only Controller under the given parent (usually a Thread Group). Its children run only on each thread's first iteration and are skipped on every later one - use it for a login before a looping scenario (or in a Thread Group that loops forever / runs on a scheduler, as find_breaking_point does).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Once Only Controller | |
| planId | Yes | ||
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and clearly reveals the key behavioral trait: children run only on the first iteration and are skipped later. It also notes the typical parent type. It does not cover side effects like plan modification or errors, but the primary runtime behavior is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is two sentences with no filler: the first states the action and scope, the second explains behavior and a concrete use case. The most important information is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 'add element' tool, the description covers purpose, placement, behavior, and a typical use case. Gaps remain around planId/name semantics and response behavior, but given the tool's low complexity, this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only vaguely refers to 'the given parent' for parentId. It does not explain planId or name at all, and the default name is only implied by the schema default. This is insufficient to fully clarify parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Add a Once Only Controller') and identifies the target parent location ('under the given parent'), immediately distinguishing it from sibling add_* tools. It further explains the core behavior of the added element, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit usage context is given: 'use it for a login before a looping scenario' and mention of the find_breaking_point scenario. However, it does not name alternative tools or state when not to use it, so it falls short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_random_controllerA
Add a Random Controller under the given parent (usually a Thread Group). On each pass it runs exactly one randomly chosen child (added under it) instead of all of them.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Random Controller | |
| planId | Yes | ||
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It successfully explains the core random-selection behavior ('runs exactly one randomly chosen child'), which is valuable. However, it does not disclose potential side effects such as the controller being created empty or any restrictions on child types, leaving some behavioral aspects unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main action and behavior, with no fluff. Every word contributes to understanding the tool, and it is appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple creation tool, the description covers purpose and behavior but lacks explicit parameter explanations, prerequisites (e.g., parent must exist and be a valid container), and what the resulting controller contains (e.g., empty children initially). It also does not mention return values or errors, leaving some gaps for an agent to handle.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining each parameter. It only implicitly references 'parent' for parentId but does not elaborate on the 'name' parameter or explicitly tie 'planId' to a context. No added meaning beyond the schema's field names, leaving agents to infer the purpose of most parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Add' and the resource 'Random Controller', and explicitly distinguishes its behavior (runs one randomly chosen child per pass) from other controllers. It is unambiguous about what the tool accomplishes and differentiates it from siblings like add_loop_controller or add_if_controller.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context by noting the parent is 'usually a Thread Group', implying where to attach it, but it does not explicitly state when to choose this controller over alternatives (e.g., when random selection is desired) or when not to use it. No direct comparison with sibling controllers is made, so usage guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_regex_extractorA
Add a Regular Expression Extractor under an HTTP sampler, to save a value from its response (body, headers, etc.) into a variable using a regex capture group. Use this instead of add_json_extractor for non-JSON (HTML/XML/plain-text) responses.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Regular Expression Extractor | |
| regex | Yes | Regular expression with at least one capture group | |
| planId | Yes | ||
| parentId | Yes | Id of the HTTP sampler node this extractor applies to | |
| template | No | Which capture group(s) to store, e.g. '$1$' | $1$ |
| matchNumber | No | Which match to use (1 = first); 0 = random match; negative = all matches | |
| defaultValue | No | Value to use if the regex doesn't match | NOT_FOUND |
| referenceName | Yes | JMeter variable name to store the extracted value in |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries the full burden of behavioral disclosure. It mentions what the tool does but not any side effects, failure conditions (e.g., if parentId is not an HTTP sampler), or whether the operation can overwrite existing elements. For an add operation that modifies the test plan, this lack of detail is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core purpose is front-loaded, and the usage guidance is efficiently integrated. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with a 75% schema coverage and no output schema, the description covers purpose and usage guidance but omits details about error handling and behavioral effects on the test plan. It is reasonably complete for a typical agent, but the lack of behavioral disclosure (as noted) prevents a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so most parameters are already documented in the schema. The description adds no new parameter-specific information beyond restating the purpose. Baseline 3 is appropriate since the schema handles the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (Add a Regular Expression Extractor), the target (under an HTTP sampler), and the purpose (save a value from response into a variable using a regex capture group). It also explicitly contrasts with add_json_extractor, distinguishing it from a sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance: 'Use this instead of add_json_extractor for non-JSON (HTML/XML/plain-text) responses.' This names the alternative and the condition that selects this tool, leaving no ambiguity about when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_response_assertionC
Add a Response Assertion under an HTTP sampler, to fail the sample if the response doesn't match.
| Name | Required | Description | Default |
|---|---|---|---|
| not | No | Negate the match (assert the pattern does NOT match) | |
| name | No | Response Assertion | |
| planId | Yes | ||
| parentId | Yes | ||
| patterns | Yes | ||
| matchType | No | contains | |
| testField | No | response_data |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the tool fails the sample when the response doesn't match, which is the primary side effect. However, it does not mention constraints (e.g., parent must be an HTTP sampler), error conditions, or any side effects on existing assertions. The core behavior is clear but details are missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the action and purpose without unnecessary words. It is concise and efficient, scoring well for conciseness, though it could be expanded slightly with parameter guidance without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, no output schema, and low schema description coverage, the description is inadequate. It doesn't explain the required parameters (planId, parentId, patterns) or optional ones (matchType, testField, name, not), nor does it describe the return behavior. An agent would need to infer too much from the schema alone, making the description incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 14% (only 'not' has a description), so the description must compensate. It does not explain planId, parentId, patterns, matchType, testField, or name. The phrase 'if the response doesn't match' implies patterns are used for matching but offers no parameter semantics. The agent is left to guess from property names and enums, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: adding a Response Assertion under an HTTP sampler, with the purpose of failing the sample on mismatch. This is specific and distinguishes it from sibling add_* tools (e.g., add_http_sampler, add_json_extractor) by naming the resource and behavior. However, it doesn't explicitly contrast with alternatives, so a 4 is appropriate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus others, nor any conditions where it should not be used. It only states what the tool does without context on selection. No exclusions or alternatives are mentioned, leaving the agent to infer appropriateness from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_setup_thread_groupA
Add a setUp Thread Group under the TestPlan root. Runs once, before all normal Thread Groups start, regardless of where it sits in the tree - typically used for one-time setup (e.g. login, seeding data).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | setUp Thread Group | |
| loops | No | Number of loop iterations per thread, -1 for infinite | |
| planId | Yes | ||
| parentId | Yes | ||
| numThreads | Yes | ||
| delaySeconds | No | Startup delay: seconds to wait before this thread group starts (enables the scheduler). Use it to stage thread groups, e.g. a second group that begins 5s after the first. | |
| durationSeconds | No | If set, run on a scheduler for this many seconds instead of a fixed loop count | |
| rampTimeSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the key execution trait: the group runs once before all normal Thread Groups regardless of tree position. However, it does not explain call-level effects such as whether the test plan is mutated immediately, what happens on invalid parentId, or what response to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the core purpose and adds a typical use case and execution nuance, so every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for selecting the tool and understanding the special setUp Thread Group behavior. However, with 8 parameters, no output schema, and no annotations, it leaves the agent to infer parameter roles and error/return behavior, making it not fully self-sufficient for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 38% and there are no annotations, so the description must compensate for undocumented parameters. It adds only the placement hint that the group is added under the TestPlan root, which partially informs parentId. It does not clarify required parameters such as planId, numThreads, or rampTimeSeconds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Add a setUp Thread Group under the TestPlan root.' It clearly distinguishes this from a normal thread group by explaining the setUp thread group's special execution behavior, so an agent can tell it apart from add_thread_group.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use case: 'typically used for one-time setup (e.g. login, seeding data).' It also clarifies that it runs before all normal Thread Groups, which helps an agent choose this tool over add_thread_group, though it does not explicitly mention the teardown sibling or say when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_size_assertionA
Add a Size Assertion under an HTTP sampler, comparing a response field's byte size against a threshold.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Size Assertion | |
| size | Yes | Size in bytes to compare against | |
| planId | Yes | ||
| operator | No | equal | |
| parentId | Yes | ||
| testField | No | response_network_size |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It states the primary effect—adding a Size Assertion and comparing byte size—but does not mention side effects, validation behavior, or what happens if the parent element is invalid. Acceptable for a simple add operation, but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the action and target, with zero filler. It is efficiently structured and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters, no output schema, and no annotations, one sentence is not enough. It lacks explanation of required IDs, operator semantics, and testField options, and provides no guidance among the many sibling add_* tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17%, so the description must compensate. It adds useful meaning for parentId ('under an HTTP sampler'), testField ('response field'), and size ('byte size'), but leaves planId, operator, and name semantics unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Add'), names the exact resource ('Size Assertion'), and states the precise purpose: comparing a response field's byte size against a threshold. This clearly distinguishes it from sibling assertion tools like add_json_assertion, add_duration_assertion, and add_response_assertion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives placement context by saying the assertion goes 'under an HTTP sampler', which is useful. However, it does not say when to choose a size assertion over the other assertion tools, nor mention any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_summary_report_listenerB
Add a Summary Report listener under the given parent (Thread Group or TestPlan). Its output is what get_execution_report reads after a run.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Summary Report | |
| planId | Yes | ||
| filename | No | Filename kept in the .jmx for portability; ignored at execution time | |
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure. It does mention that the listener's output is what get_execution_report reads, which is useful context, but it does not disclose any side effects, permissions, or reversibility of adding the listener. As a mutation tool, this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The primary action is front-loaded, and the follow-up sentence adds valuable context about the tool's relationship to get_execution_report. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters and no annotations or output schema, the description is too thin. It clarifies the parent types but omits details on planId and the return value. The link to get_execution_report is helpful, but the description does not adequately equip an agent to invoke the tool correctly with all parameters understood.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, so the description must compensate for the undocumented parameters. It provides some context for parentId by naming allowed parent types, but planId and name are not explained. The description does not clarify the meaning or format of the parameters beyond what is already in the schema, which is minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: adds a Summary Report listener under a parent (Thread Group or TestPlan). It also connects it to get_execution_report, which helps place its role. It does not explicitly contrast with the sibling add_aggregate_report_listener, so a 4 rather than 5 is appropriate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Some usage guidance is implied by mentioning the valid parents (Thread Group or TestPlan), but the description does not state when to choose this listener over the aggregate report alternative, nor any prerequisites or exclusions. Sibling tools exist but no comparison is made, so guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_tcp_samplerB
Add a TCP Sampler under the given parent (usually a Thread Group), opening a raw TCP connection and sending request data.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | TCP Sampler | |
| port | Yes | ||
| planId | Yes | ||
| server | Yes | ||
| noDelay | No | ||
| request | Yes | Data to send over the connection | |
| parentId | Yes | ||
| classname | No | TCP client handler class (short name resolves under org.apache.jmeter.protocol.tcp.sampler) | TCPClientImpl |
| timeoutMs | No | ||
| closeConnection | No | ||
| reUseConnection | No | ||
| connectTimeoutMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries full responsibility. It states it opens a connection and sends data, which is basic behavior, but it omits critical details like connection lifecycle (close/reuse), timeouts, response handling, or side effects. An agent would be unaware of several configurable behaviors that affect execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with zero filler. It efficiently conveys the core purpose without redundancy, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 12 parameters, no output schema, and no annotations, this description is far from complete. It fails to clarify connection management, timeout behavior, or what happens after sending data. An agent would need to guess at many operational aspects, making it insufficient for a complex sampler tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% (only 'request' and 'classname' have descriptions). The description adds no parameter-specific explanations beyond the schema, failing to compensate for the low coverage. Parameters like noDelay, closeConnection, reUseConnection, and timeouts are left entirely to the agent's inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (add a TCP Sampler), the resource (under a given parent, typically a Thread Group), and the specific behavior (opening a raw TCP connection and sending request data). This distinguishes it from HTTP, FTP, and other sampler siblings without needing to inspect the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides placement guidance (under a parent, usually a Thread Group) but gives no explicit when-to-use versus alternatives, no exclusions, and no mention of when not to use it. The purpose implies raw TCP usage, but the description does not actively route the agent away from other samplers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_teardown_thread_groupA
Add a tearDown Thread Group under the TestPlan root. Runs once, after all normal Thread Groups finish, regardless of where it sits in the tree - typically used for one-time cleanup.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | tearDown Thread Group | |
| loops | No | Number of loop iterations per thread, -1 for infinite | |
| planId | Yes | ||
| parentId | Yes | ||
| numThreads | Yes | ||
| delaySeconds | No | Startup delay: seconds to wait before this thread group starts (enables the scheduler). Use it to stage thread groups, e.g. a second group that begins 5s after the first. | |
| durationSeconds | No | If set, run on a scheduler for this many seconds instead of a fixed loop count | |
| rampTimeSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses genuinely valuable execution semantics: the group runs once, executes after all normal Thread Groups, and is position-independent. This adds real behavioral context beyond the tool name. It could go further (e.g., noting whether it runs on test failure), but what's disclosed is meaningful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a tight two-sentence construction with the action front-loaded. Every clause earns its place - the execution timing and cleanup purpose are both useful. It could have named the sibling for contrast, but as written it is efficient and well-organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must carry more weight. It covers behavioral semantics well but omits prerequisites (e.g., that a TestPlan root must exist and planId must reference it) and gives no guidance on the four required parameters. For an 8-parameter tool this is adequate but leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low at 38%, and the description adds nothing about parameters. planId, parentId, numThreads, rampTimeSeconds (all required) lack schema descriptions and are not explained in the tool description either. The description should compensate for the coverage gap but does not, though the IDs are somewhat self-evident from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource+location ('Add a tearDown Thread Group under the TestPlan root') and immediately distinguishes it from normal Thread Groups by stating it runs after all of them finish regardless of tree position. This clearly separates it from the sibling tools add_thread_group and add_setup_thread_group.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the execution semantics (runs once, after normal groups, position-independent) and the typical use case ('one-time cleanup'), giving clear context for when to use it. However, it doesn't explicitly name the sibling alternatives add_thread_group or add_setup_thread_group, nor state when NOT to use it, so the routing is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_thread_groupC
Add a Thread Group (virtual users) under the given parent node (usually the TestPlan root).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| loops | No | Number of loop iterations per thread, -1 for infinite | |
| planId | Yes | ||
| parentId | Yes | Id of the node to attach this thread group under | |
| numThreads | Yes | Number of concurrent virtual users | |
| delaySeconds | No | Startup delay: seconds to wait before this thread group starts (enables the scheduler). Use it to stage thread groups, e.g. a second group that begins 5s after the first. | |
| durationSeconds | No | If set, run on a scheduler for this many seconds instead of a fixed loop count | |
| rampTimeSeconds | Yes | Seconds to reach full thread count |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. The verb 'Add' implies a mutation, but the description does not state side effects, whether the operation is reversible, required permissions, or what the return value looks like. It adds no behavioral context beyond the obvious action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the core purpose and a key detail about the parent node. However, it is so terse that it lacks depth, but conciseness is a strength here.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters and no output schema, this description is inadequate. It does not explain the role of a Thread Group in a test plan, typical usage patterns, or how this tool relates to its setup/teardown siblings. The agent is left to infer too much, relying solely on the schema for semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so most parameters are already documented. The description itself adds no parameter meaning beyond what the schema provides, e.g., it does not elaborate on how numThreads or rampTimeSeconds interact. Given the moderate coverage, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Add a Thread Group') and the resource ('virtual users') and specifies the attachment point ('under the given parent node'). However, it does not distinguish this from the sibling tools add_setup_thread_group and add_teardown_thread_group, so an agent might not know which one to pick without further investigation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It only hints that the parent node is usually the TestPlan root, which is context about the parameter, not about selection criteria. No exclusions or alternative suggestions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_transaction_controllerA
Add a Transaction Controller under the given parent (usually a Thread Group). Samplers added under it (as its children) are timed and reported as a single named transaction instead of separately.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name of the transaction, e.g. 'Checkout Flow' | |
| planId | Yes | ||
| parentId | Yes | ||
| includeTimers | No | Whether to include timer/pre-processor delays in the transaction's reported time |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose side effects. It does reveal the key non-obvious behavior: child samplers are not reported separately but rolled up into a single transaction. It also names the usual parent type. It could add details like the effect of includeTimers or what happens with non-Thread Group parents, but the central behavioral trait is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, roughly 30 words, front-loaded with the verb and target resource. The behavioral explanation earns its place and no redundant phrases are present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The definition explains what the tool does and the grouping effect, but it leaves the required planId parameter undefined in both schema and description. It also doesn't state return behavior or error conditions, though there is no output schema. For a simple add operation the core is covered, but a required identifier being undocumented is a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: name and includeTimers have descriptions; planId and parentId are bare. The description helps only parentId ('under the given parent (usually a Thread Group)'), while planId remains undefined in both schema and description. includeTimers is documented in the schema but not reiterated, so the description adds modest value beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the action explicitly ('Add a Transaction Controller') and the target hierarchy ('under the given parent (usually a Thread Group)'). The second sentence defines its unique purpose among sibling controllers: grouping samplers into a single named transaction, so an agent can distinguish it from add_loop_controller, add_if_controller, and similar tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description communicates when to choose it: when samplers should be timed and reported as one named transaction instead of individually. It also hints at placement (under a Thread Group). It doesn't explicitly name exclusions or alternative tools, but the functional explanation gives enough context to select it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_uniform_random_timerA
Add a Uniform Random Timer under the given parent to pace requests with a randomized delay: delayMs +/- a random amount up to rangeMs, picked fresh before each sample.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Uniform Random Timer | |
| planId | Yes | ||
| delayMs | Yes | Base/minimum delay in milliseconds | |
| rangeMs | Yes | Maximum extra random delay added on top of delayMs | |
| parentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the placement ('under the given parent'), the randomized delay calculation, and that a fresh random value is picked before each sample. It does not mention error behavior or persistence, but the core runtime behavior is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with no filler; the action, target, placement, and calculation are all front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is reasonably complete but omits planId semantics and any indication of what happens if the parent is invalid. It covers the core behavior but not the full invocation context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, so the description must compensate. It clarifies delayMs and rangeMs through the formula, and 'under the given parent' hints at parentId, but planId is left unexplained. The formula also conflicts slightly with the schema's 'added on top of' wording ('+/-' vs '+').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Add'), a specific resource ('Uniform Random Timer'), and its purpose ('pace requests with a randomized delay'). It also gives the delay formula, which distinguishes it from sibling timers like add_constant_timer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when randomized pacing is desired. It does not explicitly name alternatives or state when not to use it, so the agent must infer the choice from the 'randomized delay' wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_user_defined_variablesA
Add a User Defined Variables config element under the given parent (TestPlan root for global variables, or a Thread Group to scope them there). Values are set once, before the test starts.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | User Defined Variables | |
| planId | Yes | ||
| parentId | Yes | ||
| variables | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the element is created, values are set once before execution, and parent determines scope. It does not mention override behavior, uniqueness constraints, or side effects on existing variables, leaving some room for uncertainty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences deliver the resource, placement options, and timing behavior with no filler. The most decision-relevant information (parent context) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple creation tool with no output schema, the description covers what the tool creates, where it places it, and when values are evaluated. It omits preconditions like plan existence or invalid parent handling, but these are not critical for an agent to make a correct first call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for parentId by explaining TestPlan root vs Thread Group scoping, but it does not explain planId, the variables array structure, or how the optional name behaves. This is partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Add') and names the exact resource ('User Defined Variables config element') plus the valid parent contexts. It clearly distinguishes this from add_user_parameters by emphasizing that values are set once before the test starts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear placement guidance: TestPlan root for global variables versus Thread Group for scoped ones. It implies the right time to use this tool (static, pre-test values) but does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_user_parametersA
Add a User Parameters pre-processor under a Thread Group, assigning each thread a different set of variable values (cycled by thread number), set once at thread start unless perIteration is true. Different from add_user_defined_variables, which sets the same values for every thread.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | User Parameters | |
| planId | Yes | ||
| parentId | Yes | ||
| valueSets | Yes | One array per thread/user slot; each inner array must have exactly one value per variableNames entry, in order | |
| perIteration | No | Re-apply values every loop iteration instead of once at thread start | |
| variableNames | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does so well: it discloses that values are cycled by thread number and applied once at thread start unless perIteration is true. It does not cover mutation side effects, permissions, or failure modes, but the 'Add' verb plus placement context makes the core intent clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with no filler. The core behavior is front-loaded, the perIteration nuance is included, and the sibling distinction earns its place at the end.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is solid: it defines element type, placement, data semantics, and the key alternative. It does not disclose return shape or error behavior, but those are minor for this kind of add-element operation and the description covers the essential agent decision-making context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (33%), so the description must compensate. It meaningfully explains valueSets as per-thread value sets cycled by thread number and defines perIteration's timing effect, while also implying parentId is the target Thread Group. planId and variableNames are not explicitly expanded, but their meanings are inferable from names and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Add'), a specific resource type ('User Parameters pre-processor'), and a placement target ('under a Thread Group'). It also differentiates the behavior from add_user_defined_variables, so an agent can distinguish this tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names the closest alternative, add_user_defined_variables, and explains the deciding difference: per-thread values vs. the same values for every thread. It also clarifies when perIteration changes the behavior, giving clear selection context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_view_results_tree_listenerA
Add a View Results Tree listener under the given parent (Thread Group or TestPlan). Its output is readable via get_execution_report when there's no Aggregate/Summary Report listener present. NOTE: captureFullData currently has no effect - execute_test_plan always runs JMeter with CSV output, and JMeter's CSV writer never emits response body/header columns regardless of this flag (only its XML output format can carry those); this option is a no-op until this server supports XML-format runs.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | View Results Tree | |
| planId | Yes | ||
| filename | No | Filename kept in the .jmx for portability; ignored at execution time | |
| parentId | Yes | ||
| captureFullData | No | Currently has no effect under this server's CSV-only execution model - see tool description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does this exceptionally well by detailing that captureFullData is a no-op, explaining why (CSV output never includes response body/header columns, only XML can), and noting the option will remain ineffective until server supports XML runs. This addresses the key hidden behavior. It does not mention other side effects (e.g., element ordering, duplicate handling) but the most important caveat is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then provides usage context, and finishes with the crucial no-op caveat. The NOTE is lengthy but technically necessary to explain the limitation. Each sentence earns its place, though the final sentence could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (no annotations, no output schema, low schema coverage), the description covers the essential context: what it does, how its output is accessed, and a major behavioral caveat. It could be more explicit about when to choose this over aggregate/summary listeners, but overall it is sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, so the description must compensate. It adds meaningful semantics for parentId (Thread Group or TestPlan) and captureFullData (no-op with explanation). It also clarifies filename's role (kept for portability, ignored at execution) beyond the schema. planId and name are not elaborated, but their purpose is reasonably inferable from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Add a View Results Tree listener') and the target resource ('under the given parent (Thread Group or TestPlan)'). It also differentiates from sibling listener tools by explaining how its output interacts with get_execution_report, distinguishing it from add_aggregate_report_listener and add_summary_report_listener.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies when its output is readable via get_execution_report (when no Aggregate/Summary Report listener is present), which provides clear context for choosing this listener. However, it does not explicitly state when to prefer this over aggregate/summary listeners, only implies it via the output-readability condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_while_controllerA
Add a While Controller under the given parent (usually a Thread Group). Samplers added under it (as its children) repeat while condition holds. Leave condition blank (or 'LAST') to loop while the last sampler in the loop succeeded; otherwise it's a JMeter expression re-evaluated each pass, looping until it evaluates to the literal string 'false'.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | While Controller | |
| planId | Yes | ||
| parentId | Yes | ||
| condition | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full behavioral burden. It transparently explains that samplers repeat while the condition holds, and details the condition semantics: blank or 'LAST' loops on last sampler success, otherwise a JMeter expression re-evaluated until it evaluates to 'false'. This is substantial and useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero fluff. It front-loads the primary action, then explains the condition logic efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavioral aspects: parent placement, looping behavior, and condition semantics. It omits return value details, but there is no output schema to contradict. It also does not mention potential error cases, but for a controller-adding tool this is acceptable. Overall it is fairly complete for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all parameters. It thoroughly explains the 'condition' parameter semantics, which is the most complex. However, it does not explicitly define planId, parentId, or name, though their meanings are somewhat inferable from the tool name and context. Given the 0% coverage, this is a partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool adds a While Controller under a parent, with a specific verb and resource. It explains the controller's purpose but does not explicitly differentiate from sibling tools like add_loop_controller or add_if_controller, which could be ambiguous to an agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives context that the parent is usually a Thread Group, which helps placement, but it does not explicitly state when to use this over alternatives like loop or if controllers. The looping semantics are described but not the decision criteria for choosing this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_xpath_extractorA
Add an XPath Extractor under an HTTP sampler, to save a value from an XML/HTML response into a variable using an XPath expression. Set tolerant=true for real-world (non-strict) HTML.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | XPath Extractor | |
| planId | Yes | ||
| parentId | Yes | Id of the HTTP sampler node this extractor applies to | |
| tolerant | No | Use a lenient HTML parser instead of a strict XML parser | |
| xpathQuery | Yes | ||
| matchNumber | No | Which match to use (1 = first) | |
| defaultValue | No | NOT_FOUND | |
| referenceName | Yes | JMeter variable name to store the extracted value in |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It provides a useful tip about setting tolerant=true for non-strict HTML, implying the parser behavior. However, it does not disclose what happens when no match occurs (defaultValue defaults to NOT_FOUND) or that matchNumber defaults to the first match, which are relevant behaviors for an extractor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, with the primary purpose front-loaded and a specific operational tip in the second sentence. There is no filler, repetition, or unnecessary detail; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple add-element operation with no output schema, the description covers the core purpose, placement, and a key behavioral toggle. Minor gaps remain around no-match handling and match selection, but the schema covers those defaults, making the overall description adequately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, so the description must compensate for the four undocumented parameters. It adds context for xpathQuery (it is an XPath expression) and clarifies the tolerant parameter's purpose, but does not explain defaultValue, name, or planId beyond what the schema already provides. Moderate value, not full compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Add'), a specific resource ('an XPath Extractor'), and its placement ('under an HTTP sampler'). It further specifies the function: extracting a value from XML/HTML responses via XPath. This clearly distinguishes it from sibling tools like add_regex_extractor and add_json_extractor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit context for when to use this tool: when the response is XML/HTML and the extraction method is XPath, placed under an HTTP sampler. It does not name alternatives or exclusions, but the conditions are specific enough for an agent to select correctly among similar add_* extractor tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_test_planA
Create a new JMeter test plan. Returns the planId and the id of its root TestPlan node, which you'll use as parentId for the first thread group.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Human-readable name for the test plan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It discloses the return values (planId and root node id) and hints at their usage, but does not mention side effects, persistence, or any prerequisites. It is adequate but not rich enough for full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero redundancy. It front-loads the core action, then concisely provides the critical return information and usage hint. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple create operation with a single parameter and no output schema, the description explains the return values and their practical use. It is complete for the immediate purpose, though it could mention potential side effects or idempotency, but that does not detract significantly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema describes the only parameter 'name' with an adequate human-readable description (100% coverage). The tool description adds no extra detail about the parameter, so it stays at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Create a new JMeter test plan' with a specific verb and resource. It further differentiates from siblings by noting that it returns planId and the root TestPlan node id, which is unique to this tool among the listed siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on how to use the return value ('which you'll use as parentId for the first thread group'), implying it is the first step in building a plan. It does not explicitly contrast with alternatives like add_thread_group, but the usage context is evident and sufficient for a create-first tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_test_planA
Start running a test plan with JMeter in non-GUI mode. Returns immediately with an executionId; the run continues in the background. Poll get_execution_status to know when it's done, then call get_execution_report to read the results.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries full responsibility. It discloses the asynchronous execution model, the immediate return of an executionId, and the necessity to poll status and retrieve results. This is critical behavioral information that an agent needs to know to avoid blocking or assuming synchronous completion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action. The second sentence packs the async behavior and follow-up steps without waste. Every sentence adds value; there is no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter asynchronous tool with no output schema, the description covers the essential operational model: what it starts, how it returns, and the required polling/report steps. It leaves no critical ambiguity for an agent to execute correctly. The mention of non-GUI mode adds practical context without going overboard.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter, planId, with 0% schema description coverage. The description does not directly explain the parameter, but its name and the tool's reference to 'test plan' make it self-explanatory as the ID of the plan to run. While the description does not add extra meaning, the parameter's purpose is evident from context. Given the low coverage, a 3 is appropriate – it does not harm, but could be more explicit about where to get planId (e.g., from list_test_plans).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Start running a test plan'), the resource ('a test plan'), the mode (non-GUI), and the immediate return with an executionId. It distinguishes itself from siblings by focusing on the initiation of execution, while get_execution_status and get_execution_report are for monitoring/results, and create_test_plan is for creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs the agent on the follow-up workflow: 'Poll get_execution_status to know when it's done, then call get_execution_report to read the results.' This tells when to use this tool (to start) and what to use next. It also states the async nature ('returns immediately... continues in the background'), providing clear context for correct usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_breaking_pointA
Find the concurrency level where a test plan stops meeting its SLA. Runs the plan over and over on its own, driving one thread group's thread count: first doubling the load until the SLA breaks, then binary-searching between the last healthy level and the first broken one. Returns immediately with a searchId; poll get_breaking_point_status for progress and the final breaking point. The thread group is temporarily switched to ramp-up + fixed-plateau (scheduler) mode for the search and its original settings are restored when the search ends. Requires an Aggregate Report, Summary Report, or View Results Tree listener in the plan. Rounds always run the thread group in scheduler mode with an infinite loop count, so every thread repeats its whole scenario until the plateau ends: a request meant to run once per user (e.g. a login) must sit under a Once Only Controller, and a nested Loop Controller runs its count on every repeat. Each round reports its metrics both overall and per label (byLabel, with each label's share of the samples) - check that the mix matches what the plan intends. The SLA is judged on the overall numbers, not per label.
| Name | Required | Description | Default |
|---|---|---|---|
| p95Ms | No | Fail a round when overall p95 latency exceeds this many milliseconds. | |
| planId | Yes | ||
| errorPct | No | Fail a round when the overall error rate exceeds this percentage. | |
| maxThreads | Yes | Safety ceiling - the search never runs more threads than this. | |
| startThreads | No | Thread count for the first round; doubles from there while the SLA holds (default 50). | |
| maxIterations | No | Hard cap on rounds, so a slow search still ends (default 8). | |
| cooldownSeconds | No | Pause between rounds so the system under test recovers (default 5). | |
| toleranceThreads | No | Stop bisecting once the healthy and broken bounds are this close (default: 2% of maxThreads, minimum 5). | |
| threadGroupNodeId | Yes | Node id of the thread group whose thread count the search will drive. | |
| rampSecondsPerThread | No | Ramp-up seconds per thread, so every round adds load at the same rate (default 0.1). | |
| plateauDurationSeconds | No | Seconds to hold full load after ramp-up; only these samples count toward the SLA (default 60). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it delivers extensively. It reveals that the plan runs repeatedly, the thread group is temporarily switched to scheduler mode and restored afterward, rounds use an infinite loop count, and SLA judgment is based on overall metrics rather than per-label. These are non-obvious side effects and semantics that an agent must know before invoking the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: purpose, algorithm, return behavior, mode-switching side effect, listener prerequisite, loop-count caveat, and metric semantics. It is front-loaded with the core purpose and progressively adds necessary operational detail without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema, the description is remarkably complete. It explains the asynchronous return pattern, the required listener, the thread-group mutation and restoration, the per-label vs overall metric distinction, and the SLA evaluation basis. An agent has enough information to invoke the tool correctly and interpret its behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 91%, so the schema already documents most parameters well. The description adds useful context about how parameters relate to the search algorithm (e.g., doubling, binary search, overall SLA), but it does not substantially elaborate individual parameter meanings beyond what the schema provides. This matches the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Find the concurrency level where a test plan stops meeting its SLA.' It then details the exact algorithm (doubling then binary search), which distinguishes it clearly from sibling execution and status tools. The purpose is unambiguous and not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational guidance: it returns a searchId to poll via get_breaking_point_status, and it states the prerequisite that an Aggregate Report, Summary Report, or View Results Tree listener must exist. It does not explicitly contrast with execute_test_plan or normal test runs, but the specialized search behavior and polling instructions make the intended usage evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_breaking_point_statusA
Check a capacity search started with find_breaking_point: status (running/completed/failed/stopped), progress (rounds completed out of maxIterations, plus the in-flight round's elapsed time and percent complete), the rounds run so far with each one's load and overall + per-label metrics, and - once it finishes - the breaking point, the last healthy load, and a plain-language conclusion. breakingPoint is the lowest load that was tested and broke the SLA, not necessarily the exact edge: read breakingPointRange (healthyUpTo / brokenAt / exact) for the real precision, since levels between the two were never run when toleranceThreads is above 1. There is no completion push - poll this tool. The same data is on disk at files.meta, and each round's raw results are in files.executionsDir//.
| Name | Required | Description | Default |
|---|---|---|---|
| searchId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the lack of push notifications, explains the precision limitation of breakingPoint (reading breakingPointRange for exactness), and mentions where data is persisted on disk. It does not explicitly state the operation is read-only, but 'Check' implies it, and the detailed result semantics add value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense, front-loading the core purpose and then detailing nuances. Each sentence adds value—progress details, precision caveat, polling behavior, and disk locations. It is structured logically and not overly verbose for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description thoroughly covers what the tool returns: status, progress, rounds with metrics, breaking point, conclusion, and breakingPointRange. It also explains the precision caveat and where raw data lives. Missing edge cases (e.g., invalid searchId) are minor and do not detract from overall completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It implies that searchId is the identifier returned by find_breaking_point by referencing that tool, which gives context. However, it does not explicitly state the origin or format of searchId, leaving room for slight ambiguity. For a single string parameter, this is adequate but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Check' and the resource 'capacity search started with find_breaking_point', and enumerates exactly what is returned (status, progress, rounds, breaking point, conclusion). It implicitly distinguishes from sibling tools like get_execution_status by specifying the breaking-point search context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates this tool is for checking a search started with find_breaking_point and explicitly notes 'There is no completion push - poll this tool', which guides usage for polling. However, it does not name alternative tools for other execution status checks or explicitly state when not to use it, leaving that to inference from siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_execution_reportA
Read and aggregate the results of a finished (or still-running) execution, computed from its Aggregate Report / Summary Report listener output: per-label and overall count, error rate, avg/min/max/median/p90/p95/p99 latency, throughput and KB/sec.
| Name | Required | Description | Default |
|---|---|---|---|
| executionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the operation is read-only ('Read'), the data source (listener output), and that it can be called on running executions. It does not mention potential side effects, latency, or error handling for invalid IDs, but for a read operation this is acceptable. It adds context about the metrics but not deeper behavioral details like caching or consistency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence that front-loads the core purpose and immediately lists the output metrics. Every word contributes value; there is no fluff or repetition. It is concise without losing specificity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers the essential elements: what it does, when it can be used (finished or running), and exactly which metrics it returns. The lack of an output schema is mitigated by the metrics list. It does not mention error conditions or return formatting, but that is a minor gap given the simplicity of the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has a single parameter, executionId, with 0% description coverage. The tool description does not explicitly explain what the parameter means beyond its name—it only says the tool retrieves results for 'an execution', but does not reiterate that executionId must be the ID of the execution whose report is desired. Since schema coverage is zero and the description fails to compensate, the parameter semantics are under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read and aggregate') with a clear resource ('results of a finished (or still-running) execution') and enumerates the exact metrics returned (per-label and overall count, error rate, latency percentiles, throughput, KB/sec). It distinguishes itself from siblings like get_execution_status by focusing on detailed report data from listener output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the tool works on finished or still-running executions and specifies the source (Aggregate Report / Summary Report listener output), giving context for when it applies. However, it does not explicitly contrast it with siblings like get_execution_status, nor does it state when NOT to use it (e.g., for just checking status). Usage is implied but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_execution_statusA
Check the status of a test run started with execute_test_plan (running/completed/failed), how long it has been running, and - while it runs - a progress block computed from the results file so far: samples, error rate, avg/p95 latency, overall and recent (last ~10s) throughput, and the same per label. Poll this tool rather than tailing stdout.log. The logTail is JMeter's console output with JVM startup warnings filtered out; JMeter only prints a summary line there about every 30 seconds, so it can stay the same between quick polls. Once the run ends, call get_execution_report for the full aggregate.
| Name | Required | Description | Default |
|---|---|---|---|
| executionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it handles it well. It discloses that progress is computed from the results file so far, that logTail is filtered JMeter console output, and that JMeter only prints a summary line about every 30 seconds so the logTail can appear unchanged between quick polls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and returned data, then compactly covers polling behavior, logTail semantics, and end-of-run routing. Every sentence earns its place; there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter polling tool with no output schema, the description is complete. It explains statuses, runtime info, the progress fields, throughput windows, per-label breakdown, logTail caveats, and the correct sibling to call after completion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for executionId (0% coverage), but the description compensates by tying it to a 'test run started with execute_test_plan', making it clear that the parameter identifies the run initiated by that tool. It adds meaning beyond the bare parameter name and required field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check the status of a test run started with execute_test_plan'. It clearly defines the possible states (running/completed/failed), what data is returned, and distinguishes itself from get_execution_report by saying the report is for after the run ends.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit guidance: 'Poll this tool rather than tailing stdout.log' and 'Once the run ends, call get_execution_report for the full aggregate'. This tells the agent exactly when to use this tool and when to switch to a sibling, which is strong usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_test_planA
Get the full element tree of a test plan, including every node's id (needed as parentId for add_* tools) and type.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It clearly indicates this is a retrieval operation ('get') and specifies what it returns, which is enough to infer it is read-only and non-destructive. It does not mention potential side effects or response size, but for a getter with a single parameter, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clearly structured sentence that front-loads the primary action ('Get the full element tree') and then adds contextual detail (ids and their use). Every word serves a purpose with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one parameter and no output schema, the description fully explains what is returned (the full element tree, including ids and types) and why it matters (for parentId). It gives an agent enough to decide when and how to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter (planId) with 0% description coverage. The description does not explicitly explain what planId is, though it is easily inferred as the identifier of the test plan from the context. It adds no direct parameter detail, but the single, self-explanatory parameter name mitigates the weak coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Get the full element tree') and a specific resource ('a test plan'), and clarifies it returns every node's id and type. This distinguishes it from siblings like list_test_plans (which lists plans) and get_execution_report (which reports execution results), making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly mentions that node ids are needed as parentId for add_* tools, giving a clear reason to use this tool before making structural modifications. It does not explicitly say when not to use it, but the context of adding nodes is clear, and the sibling tools for execution vs. structure are implicitly separated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_test_plan_xmlA
Serialize a test plan to its JMeter .jmx XML, without running JMeter.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It usefully discloses the key non-obvious trait that JMeter is not run, but it does not describe other side effects, return behavior, whether the plan must already exist, or error behavior. Some transparency is provided, but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence, front-loaded with the action and output, and includes a necessary behavioral qualifier. Every word earns its place; there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the core operation but leaves the return format implicit. It does not explicitly state that the JMX XML is returned to the caller, nor mention prerequisites or failure modes, so an agent has some residual ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and one required parameter, planId. The description does not explicitly explain planId, but the tool description's 'Serialize a test plan' combined with the parameter name makes its purpose reasonably inferable. It adds little beyond the schema, hence a mid-range score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Serialize') and resource ('a test plan'), names the exact output format ('JMeter .jmx XML'), and explicitly distinguishes this from execution with 'without running JMeter.' This clearly separates it from siblings like execute_test_plan and get_test_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without running JMeter' implies this tool is for obtaining the XML representation rather than executing a test, but it does not explicitly name alternatives or state when to prefer this over get_test_plan or execute_test_plan. Usage context is clear but exclusions are left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_test_planA
Import an externally authored .jmx file (e.g. exported from the JMeter GUI) as a new test plan. Element types this server doesn't model are kept as opaque UnknownElement nodes (their original XML is preserved and re-emitted as-is) instead of being dropped - check unknownElementCount/unknownElementTypes in the response to see what wasn't fully understood.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Plan name to use; defaults to the imported TestPlan element's name | |
| filePath | Yes | Absolute path to the .jmx file to import |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and discloses a key behavior: unmodeled element types are preserved as UnknownElement nodes with original XML re-emitted as-is, rather than dropped. It also directs the agent to response fields for unrecognized elements, adding useful context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two dense sentences with no filler. The first sentence states the purpose, and the second adds important preservation behavior and response guidance, keeping the most relevant information front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, preservation behavior, and points to response fields, which is especially useful given there is no output schema. However, it does not describe the full response shape or what happens after a successful import, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented with clear descriptions. The tool description adds file-format context but does not add meaningful parameter-level semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Import an externally authored .jmx file' with the outcome 'as a new test plan.' It is distinct from sibling add_* tools and create_test_plan because it emphasizes external authorship and file-based import.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear context for use: importing externally authored .jmx files, such as those exported from the JMeter GUI. It does not explicitly name alternatives or exclusions, but the intended trigger is evident from the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_test_plansA
List all test plans in the workspace.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states that it lists test plans, without indicating whether the operation is read-only, if there are side effects, rate limits, or pagination behavior. For a potentially safe listing operation, the lack of any safety or side-effect context is a gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that clearly states the action and scope. It is front-loaded with the verb and resource, avoiding any extraneous text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (no parameters) and lack of an output schema, the description provides adequate context for most use cases. However, it does not detail the return format (e.g., list of test plan IDs, full objects, or metadata), which could be relevant for agents expecting a specific structure. Since no output schema exists, a brief note on return content would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. No parameter documentation is needed, and the description adds nothing beyond the schema, which is appropriate since there is nothing to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb ('List') and a specific resource ('all test plans in the workspace'). It naturally distinguishes itself from sibling tools like get_test_plan (which targets a single plan) and create_test_plan (which creates), so an agent can identify its purpose without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for enumerating all test plans, but it does not explicitly state when to use it versus alternatives like get_test_plan for a specific plan. No exclusions or conditions are mentioned, so the guidance remains implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_elementA
Move an element (and its subtree) to a new parent, optionally at a specific index among the new parent's children (default: appended last). Rejects moving a node into its own subtree.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | ||
| nodeId | Yes | ||
| planId | Yes | ||
| newParentId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the entire subtree moves, that index defaults to appending last, and that moving into the node's own subtree is rejected. It does not state whether the operation is reversible or what the response contains, but the core mutation behavior is well specified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and resource, then packs the optional index behavior and rejection rule without redundancy. Every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation with no annotations and no output schema, the description covers the core move behavior and a key edge case. It omits whether moving to the current parent is allowed (relevant given reorder_children sibling), what a successful move returns, and planId's role, so an agent may still have open questions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains index (specific position, default appended last) and clarifies newParentId/nodeId roles through the sentence. However, planId is never described, leaving one required parameter's semantics to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Move') and resource ('an element (and its subtree) to a new parent'), with optional index semantics. This clearly differentiates it from siblings like reorder_children and remove_element.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or alternatives are provided. It does not mention that reorder_children should be used for same-parent reordering, nor when move_element is preferred over update_element. The only usage constraint is the self-subtree rejection, which is a validity rule rather than selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_elementA
Remove an element (and its subtree) from a test plan. The root TestPlan node cannot be removed.
| Name | Required | Description | Default |
|---|---|---|---|
| nodeId | Yes | ||
| planId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the core behavior (removes element and subtree) and a key constraint (root cannot be removed), but does not mention side effects, return values, or whether the operation is destructive beyond the obvious. It provides minimal but useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise sentence plus a constraint, with no filler. The action is front-loaded and the constraint is clearly stated, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (two parameters, no output schema), but the description leaves parameter semantics unexplained and does not describe what happens after removal (e.g., whether the plan is returned). It is adequate for a basic understanding but incomplete for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not mention nodeId or planId at all. The agent must infer which parameter refers to the element and which to the plan. This is a significant gap for a tool with two required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Remove') and resource ('an element (and its subtree) from a test plan'), clearly distinguishing it from sibling tools like rename_element, update_element, and move_element. The added constraint about the root TestPlan node further clarifies scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for deleting elements but does not explicitly mention when to use it versus alternatives, nor does it provide any exclusions beyond the root node. There is no mention of when to prefer this over other mutation tools, so guidance is only implied by the action itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_elementC
Rename an element (its testname in the generated .jmx).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| nodeId | Yes | ||
| planId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It only mentions that the rename affects the testname in the .jmx, giving a small hint, but it does not disclose side effects, reversibility, permission requirements, or what happens to existing references. For a mutation tool, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, direct sentence that front-loads the action. It is appropriately concise with no filler, though its brevity comes at the cost of missing essential context, which slightly reduces the score from a perfect 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 required parameters, no annotations, no output schema, and a 0% schema description coverage, the description is severely inadequate. It fails to explain any parameter semantics, usage context, or behavioral consequences, making it nearly impossible for an agent to call it correctly without external knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for explaining the three required parameters (planId, nodeId, name). It does not mention any of them, leaving the agent to guess which parameter is the new name and what the others mean. This is a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (rename) and the resource (element), and adds specific context that it affects the testname in the generated .jmx. It distinguishes from siblings like update_element or remove_element by focusing on the rename operation, though it doesn't explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: if you need to rename an element, use this tool. However, it provides no explicit conditions, prerequisites, or comparisons with sibling tools like update_element, leaving the agent to infer when this is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reorder_childrenA
Reorder a node's direct children. orderedChildIds must be an exact permutation of that node's current children ids.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| parentId | Yes | ||
| orderedChildIds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the critical precondition that orderedChildIds must be an exact permutation of the node's current child ids, which is essential for correct invocation. It does not describe failure behavior, but the key constraint is surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences under 20 words total, with the operation front-loaded and the constraint following immediately. There is no filler, repetition, or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter reorder operation with no output schema, the description provides the core operation and the essential constraint. It does not mention return value or how to obtain current child ids, but these are minor given the tool's simplicity and clear parameter naming.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only elaborates on orderedChildIds (stating it must be an exact permutation), leaving planId and parentId to be inferred solely from their names. It adds some semantic value for the key parameter but does not fully cover all three.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Reorder') and a precise resource ('a node's direct children'), immediately distinguishing it from sibling tools like move_element or the add_* family. The phrase 'direct children' clearly scopes the operation, and there is no tautology or ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (reordering a node's direct children) but does not explicitly state when to prefer this tool over siblings such as move_element or update_element. No alternatives or exclusions are provided, so guidance is only implied, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_element_enabledA
Enable or disable an element without removing it. A disabled element is skipped by JMeter at run time.
| Name | Required | Description | Default |
|---|---|---|---|
| nodeId | Yes | ||
| planId | Yes | ||
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It clearly states that disabling does not remove the element and that disabled elements are skipped by JMeter at run time. While it does not detail persistence or reversibility, these are reasonably implied by 'enable or disable', and the essential runtime behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences deliver the core action, the non-destructive nature, and the runtime behavior with no filler. Every word adds value and the most important distinction ('without removing it') is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter toggle, the description covers the primary behavior and runtime effect, but the 0% schema coverage leaves the identification of planId and nodeId entirely to naming conventions. The lack of annotations and absence of an output schema also means the description alone must answer all usage questions, and it does so adequately but not exhaustively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. It provides semantic context for the 'enabled' parameter by explaining what a disabled element does at runtime, but it does not explain 'planId' or 'nodeId' beyond their self-evident names, leaving required parameters only partially clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb pair ('Enable or disable') with a clear resource ('an element') and explicitly contrasts with removal, which differentiates it from sibling tools like remove_element. It also adds the JMeter-specific consequence, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without removing it' gives clear context for when this tool is appropriate: when the element should remain in the test plan but be skipped at runtime. It does not explicitly name alternatives like remove_element, but the implied usage boundary is clear and easily actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_breaking_point_searchA
Stop a running capacity search. Terminates the round in flight, restores the thread group's original settings, and keeps whatever bounds the search had established so far.
| Name | Required | Description | Default |
|---|---|---|---|
| searchId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses key behavioral effects: terminates the round, restores original thread group settings, and retains established bounds. However, it does not mention what happens if the search is not running, whether the operation is idempotent, or any side effects on related resources. The disclosure is useful but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the primary action ('Stop a running capacity search') and then lists the three key behavioral outcomes. There is no fluff, and every word contributes to the tool's meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool, the description covers the essential behavior: what it stops, what it restores, and what it preserves. It does not describe the return value (no output schema) or error handling, but these are not critical for an agent to invoke it correctly. The lack of explicit conditions (e.g., 'only call when a search is active') is a minor gap, but the phrase 'running capacity search' already implies this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain the searchId parameter. While the parameter name is self-explanatory, the description should ideally state that searchId is the identifier returned by find_breaking_point or obtained from a status call. Since the description adds no semantic value beyond the schema, it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Stop' and the resource 'capacity search', and distinguishes it from siblings like find_breaking_point (which starts a search) and get_breaking_point_status (which queries status). The additional detail about terminating the round, restoring settings, and preserving bounds makes the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to halt an in-progress search, but it does not explicitly state when to use it versus alternatives or mention any prerequisites (e.g., a search must be running). There is no guidance on error conditions or fallback tools. The naming and context make the usage reasonably clear, but explicit routing is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_executionA
Stop a running test execution (sends SIGTERM to the JMeter process).
| Name | Required | Description | Default |
|---|---|---|---|
| executionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavior. It does reveal the internal action (SIGTERM), but omits consequences such as whether the stop is graceful, what happens if the execution is already stopped, or whether any cleanup occurs. It also does not mention potential side effects or error states. The provided detail is useful but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tight sentence with no filler. It front-loads the primary action and includes the key behavioral detail (SIGTERM) without redundancy. Every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema. The description covers the core action and mechanism, but given the absence of annotations and output schema, it could state what the caller should expect (e.g., a success status or error if execution not found). It also omits edge cases like calling on an already stopped execution. These gaps are moderate for a low-complexity tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'running test execution' which loosely implies the executionId identifies that execution, but it does not explicitly explain the parameter. For a single unambiguous identifier, this minimal inference is acceptable, but the description adds no explicit clarification beyond the schema's bare type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Stop a running test execution') and the specific mechanism ('sends SIGTERM to the JMeter process'). It distinguishes itself from siblings like execute_test_plan (which starts) and get_execution_status (which reports status) by focusing on termination. The verb and resource are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it should be used on a running execution but does not explicitly state when to use it versus alternatives. It lacks exclusions, such as 'only use if execution is in a running state' or 'for status checks, use get_execution_status instead'. No guidance is provided on prerequisite conditions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_elementA
Update an element's props with a shallow merge (or full replace). A prop value of null in props removes that key. When the node's type is one of the modeled types, the resulting props are validated against that type's schema (check the type first via get_test_plan).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | merge | |
| props | Yes | Patch to apply to the node's props | |
| nodeId | Yes | ||
| planId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses shallow merge vs. replace, null-key removal, and validation for modeled types. However, it omits error handling on validation failure, permission requirements, or reversibility, leaving notable behavioral gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action (shallow merge/full replace). The second sentence adds an important validation caveat without excess. No fluff, but it could be slightly more structured to separate the validation note from the core behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description covers the main update semantics and the validation caveat. However, it does not mention what happens if validation fails, whether the operation is reversible, or any prerequisites beyond checking the type. These omissions leave the context incomplete for safe and effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only 'props' has a description). The description adds significant meaning by explaining shallow merge semantics, null removal, and validation implications for props and mode. It does not elaborate on planId and nodeId, but those are self-explanatory identifiers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (update) and resource (element's props) and specifies the merge/replace modes. It distinguishes itself from sibling add_* and rename_element tools by focusing on modifying existing element props, so an agent can tell it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a precondition ('check the type first via get_test_plan') for validation, which is useful context. However, it does not explicitly state when to use this tool versus alternatives (e.g., adding a new element), so usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
44 tool updates
v0.4.5- Added
add_backend_listener - Added
add_constant_throughput_timer - Added
add_constant_timer - Added
add_cookie_manager - Added
add_csv_data_set - Added
add_duration_assertion - Added
add_ftp_request - Added
add_http_request_defaults - Changed
add_http_sampler5 fields changed- changed
Input schema / properties / domain / descriptionPrevious value: -"Host name, e.g. api.example.com"New value: +"Host name, e.g. api.example.com. Omit to inherit from HTTP Request Defaults" - added
Input schema / properties / port / descriptionAdded value: +"Omit to inherit from HTTP Request Defaults" - removed
Input schema / properties / protocol / defaultRemoved value: -"https" - added
Input schema / properties / protocol / descriptionAdded value: +"Omit to inherit from HTTP Request Defaults" - changed
Input schema / requiredPrevious value: -[ - "planId", - "parentId", - "name", - "method", - "domain", - "path" -]New value: +[ + "planId", + "parentId", + "name", + "method", + "path" +]
- Added
add_if_controller - Added
add_interleave_controller - Added
add_jdbc_connection_configuration - Added
add_jdbc_request - Added
add_json_assertion - Added
add_jsr223_postprocessor - Added
add_jsr223_preprocessor - Added
add_jsr223_sampler - Added
add_loop_controller - Added
add_once_only_controller - Added
add_random_controller - Added
add_regex_extractor - Added
add_setup_thread_group - Added
add_size_assertion - Added
add_tcp_sampler - Added
add_teardown_thread_group - Changed
add_thread_group1 field changed- added
Input schema / properties / delaySecondsAdded value: +{ + "description": "Startup delay: seconds to wait before this thread group starts (enables the scheduler). Use it to stage thread groups, e.g. a second group that begins 5s after the first.", + "exclusiveMinimum": 0, + "type": "number" +}
- Added
add_transaction_controller - Added
add_uniform_random_timer - Added
add_user_defined_variables - Added
add_user_parameters - Added
add_view_results_tree_listener - Added
add_while_controller - Added
add_xpath_extractor - Added
find_breaking_point - Added
get_breaking_point_status - Added
get_test_plan_xml - Added
import_test_plan - Added
move_element - Added
remove_element - Added
rename_element - Added
reorder_children - Added
set_element_enabled - Added
stop_breaking_point_search - Added
update_element
14 tool updates
v0.1.2- First observed
add_aggregate_report_listener - First observed
add_header_manager - First observed
add_http_sampler - First observed
add_json_extractor - First observed
add_response_assertion - First observed
add_summary_report_listener - First observed
add_thread_group - First observed
create_test_plan - First observed
execute_test_plan - First observed
get_execution_report - First observed
get_execution_status - First observed
get_test_plan - First observed
list_test_plans - First observed
stop_execution
TDQS
Scored across 56 tools
Each tool targets a distinct JMeter element or operation: add_* tools for different element types (thread groups, samplers, controllers, timers, assertions, extractors, listeners, config elements), and separate tools for lifecycle operations (create, update, move, remove, execute, report). Even similar tools like add_json_extractor vs add_regex_extractor are clearly differentiated by response format. There is no ambiguity in selecting between them.
All tool names follow a strict verb_noun snake_case pattern: add_* for creation, get_* for retrieval, execute/stop for control, etc. The verb is always first and lowercase, and nouns are descriptive (e.g., add_http_sampler, set_element_enabled, get_execution_report). No camelCase or mixed conventions are present.
With 56 tools, the surface is significantly larger than typical MCP servers (calibration marks 25+ as excessive). While JMeter's complexity justifies a broad set, the count still feels heavy and might overwhelm agents. Some tools could be consolidated (e.g., different extractors could be parameterized), but as-is, the tool count is beyond the 'reasonable' range.
The tool set covers the full JMeter workflow: plan creation/import, element manipulation (add, remove, update, move, enable), execution control (execute, stop, status), and result reporting (aggregate, summary, view results tree). It includes a wide variety of samplers, controllers, timers, assertions, and extractors. The main gap is the lack of a tool to delete a test plan entirely (root node cannot be removed), leaving plans to accumulate. Minor gap, but otherwise comprehensive.
Maintenance
Related MCP Connectors
Manage Timequip projects, tasks, comments, members, and dashboards through MCP.
Read and write Mission Control state via MCP — projects, tasks, subtasks, templates, status updates.
MEOK MCP Test MCP — golden-file + schema-drift + tool-failure tests for any MCP server. Drop-in
Related MCP Servers
- FlicenseCqualityDmaintenanceA Model Context Protocol server that enables executing and interacting with JMeter tests through MCP-compatible clients like Claude Desktop, Cursor, and Windsurf.2-
- FlicenseCqualityDmaintenanceA Model Context Protocol server that enables execution of JMeter performance tests through AI assistants and MCP-compatible clients like Claude, Cursor, and Windsurf.2-
- FlicenseAqualityDmaintenanceEnables the execution and analysis of JMeter performance tests through MCP-compatible clients. It provides tools for running tests in non-GUI mode, identifying performance bottlenecks, and generating comprehensive insights and visualizations from result files.6-
- FlicenseNot gradedqualityCmaintenanceIntegrates Apache JMeter with AI assistants to run and manage load tests through natural language. It enables users to execute test plans, parse results, inspect test structures, and compare performance metrics across different runs.-