Skip to main content
Glama
juliodelimas

jmeter-mcp-server

by juliodelimas
README.md
# jmeter-mcp-server

[![npm version](https://img.shields.io/npm/v/jmeter-mcp-server.svg)](https://www.npmjs.com/package/jmeter-mcp-server)
[![CI](https://github.com/juliodelimas/jmeter-mcp-server/actions/workflows/ci.yml/badge.svg)](https://github.com/juliodelimas/jmeter-mcp-server/actions/workflows/ci.yml)
[![Node.js](https://img.shields.io/node/v/jmeter-mcp-server.svg)](https://www.npmjs.com/package/jmeter-mcp-server)
[![TypeScript](https://img.shields.io/badge/TypeScript-5%2B-3178C6?logo=typescript&logoColor=white)](https://www.typescriptlang.org/)
[![License: MIT](https://img.shields.io/npm/l/jmeter-mcp-server.svg)](LICENSE)

Give an LLM real, deterministic control over [Apache JMeter](https://jmeter.apache.org/) — build test plans, run real load tests, and read back real results, without ever hand-writing `.jmx` XML or opening the GUI.

```
"Load test my API: 20 users hitting POST /orders for 2 minutes,
 5% think time, fail anything over 800ms"
```

...turns into a running JMeter test and a real report, through typed tool calls an MCP client (Claude Code, Claude Desktop, etc.) makes directly.

## Why not just ask an LLM to write the `.jmx` itself?

It can — a `.jmx` is just XML, and any capable model has seen plenty of JMeter test plans. The problem is *how* it fails: JMeter's format is a `hashTree` with dozens of fragile, easy-to-misremember details — exact `guiclass`/`testclass` pairs, property names that don't match their GUI label (`ThreadGroup.num_threads` is a `stringProp`, not an `intProp`), integer *bitmasks* for assertion match types, strict parent/child pairing with sibling `<hashTree>` tags. None of it is self-checking. A wrong value still produces valid, loadable XML that just quietly does the wrong thing.

That's not hypothetical — it happened building this project. An early version of the If Controller generated a property called `useExpression` set to `true`, which reads like "yes, evaluate my condition." The real JMeter source does the opposite: `useExpression=true` means *don't* evaluate it as an expression — just check if the string is literally `"true"`. Every non-trivial condition silently, permanently failed. No error, no warning — the child sampler just never ran. It only surfaced by actually executing the generated plan against real JMeter and noticing a sample count of zero.

That's the whole case for this server in one story: an LLM regenerating XML from memory re-risks that exact mistake on every single request. This server encodes the correct shape **once**, in a serializer (and a matching parser for the reverse direction) checked against real JMeter source and bundled examples, exercised by [166 automated tests](#testing) including real JMeter runs — and exposes it as typed tools instead. Concretely:

- **Correctness through one tested code path**, not regenerated-from-memory XML every time.
- **Cheap incremental edits.** Plans are a small JSON tree with stable node ids — adding, removing, renaming, moving, or disabling an element is one tool call by `id`, not rewriting a whole `.jmx` file. An existing `.jmx` (hand-written or exported from the GUI) can be imported and edited the same way.
- **Aggregated results, not raw samples.** `get_execution_report` returns computed stats (error %, avg/median/p90/p95/p99, throughput) — not thousands of sample rows to average by hand.
- **Real async execution.** `execute_test_plan` returns immediately with an `executionId`; long-running load tests never block anything.

The generated `.jmx` is standard JMeter output — open it in the real GUI any time.

## Example

```
You:    Build a load test: 10 users for 30s hitting GET https://api.example.com/health,
        fail anything that takes over 500ms, then run it and tell me the p95.

Claude: [create_test_plan, add_thread_group, add_http_sampler, add_duration_assertion,
         add_aggregate_report_listener, execute_test_plan, get_execution_status, get_execution_report]

        Ran 300 requests over 30s, 0 failures. p95 latency: 214ms, avg: 187ms, throughput: 10.1 req/s.
```

Every step above is a real typed MCP tool call — see [Tools](#tools) for the full set (35 element types across samplers, controllers, timers, extractors, assertions, and listeners, plus editing, inspection, and `.jmx` import/export tools) and [Example workflow](#example-workflow) for the raw call sequence.

## Quick start

```bash
claude mcp add jmeter \
  -e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
  -- npx -y jmeter-mcp-server
```

That's it — no cloning, no build. Adjust `JMETER_HOME` to your JMeter install (see [Prerequisites](#prerequisites)). Full setup details, Claude Desktop config, and local-dev instructions are in [Adding this server to Claude Code](#adding-this-server-to-claude-code).

## How a test plan is represented

Each plan is a JSON tree (`{id, type, props, children[]}`), not XML text. Authoring tools append a child under a given `parentId`, editing tools (`remove_element`, `update_element`, `move_element`, etc.) mutate that same tree in place, and the tree is only serialized to a real `.jmx` file on demand (`get_test_plan_xml`) or at execution time. `import_test_plan` runs the reverse direction, parsing an existing `.jmx` back into this same tree shape. This is what makes incremental edits cheap and keeps all the fiddly XML schema knowledge in two places (`src/jmx/serializer.ts` for tree → XML, `src/jmx/parser.ts` for XML → tree, sharing prop shapes from `src/jmx/propTypes.ts`) instead of spread across every tool.

## Tools

**Authoring** (each returns the new node's `id`, used as `parentId` for
whatever you attach under it next) — grouped the same way JMeter's own
right-click **Add** menu groups them, so if you already know the GUI, you
already know where to look:

| Tool | Adds |
|---|---|
| `create_test_plan` | Root `TestPlan` node — returns `planId` and the root node id |

**Threads (Users):**

| Tool | Adds |
|---|---|
| `add_thread_group` | Thread Group (virtual users); optional `durationSeconds` (scheduler) and `delaySeconds` (start delay, for staging groups) |
| `add_setup_thread_group` | setUp Thread Group (runs once before all Thread Groups) |
| `add_teardown_thread_group` | tearDown Thread Group (runs once after all Thread Groups) |

**Sampler:**

| Tool | Adds |
|---|---|
| `add_http_sampler` | HTTP Request sampler |
| `add_jdbc_request` | JDBC Request sampler |
| `add_jsr223_sampler` | JSR223 Sampler (Groovy/BeanShell/JS/JEXL script as the sample) |
| `add_ftp_request` | FTP Request sampler |
| `add_tcp_sampler` | TCP Sampler |

**Logic Controller:**

| Tool | Adds |
|---|---|
| `add_transaction_controller` | Transaction Controller (groups child samplers into one named transaction) |
| `add_loop_controller` | Loop Controller (repeats child samplers) |
| `add_if_controller` | If Controller (conditionally runs child samplers) |
| `add_while_controller` | While Controller (repeats children while a condition holds) |
| `add_random_controller` | Random Controller (runs one random child per pass) |
| `add_interleave_controller` | Interleave Controller (alternates through children) |
| `add_once_only_controller` | Once Only Controller (children run only on each thread's first iteration) |

**Config Element:**

| Tool | Adds |
|---|---|
| `add_csv_data_set` | CSV Data Set Config (parameterization from a file) |
| `add_user_defined_variables` | User Defined Variables |
| `add_jdbc_connection_configuration` | JDBC Connection Configuration (pooled datasource) |
| `add_http_request_defaults` | HTTP Request Defaults |
| `add_cookie_manager` | HTTP Cookie Manager |
| `add_header_manager` | HTTP Header Manager |

**Timer:**

| Tool | Adds |
|---|---|
| `add_constant_timer` | Constant Timer (pacing/think-time) |
| `add_uniform_random_timer` | Uniform Random Timer (randomized pacing) |
| `add_constant_throughput_timer` | Constant Throughput Timer (target rate pacing) |

**Pre Processors:**

| Tool | Adds |
|---|---|
| `add_jsr223_preprocessor` | JSR223 PreProcessor |
| `add_user_parameters` | User Parameters (per-thread variable value sets) |

**Post Processors:**

| Tool | Adds |
|---|---|
| `add_json_extractor` | JSON Extractor post-processor |
| `add_regex_extractor` | Regular Expression Extractor post-processor |
| `add_xpath_extractor` | XPath Extractor post-processor |
| `add_jsr223_postprocessor` | JSR223 PostProcessor |

**Assertions:**

| Tool | Adds |
|---|---|
| `add_response_assertion` | Response Assertion |
| `add_json_assertion` | JSON Assertion (JSONPath validation) |
| `add_duration_assertion` | Duration Assertion (response-time SLA) |
| `add_size_assertion` | Size Assertion (response byte-size check) |

**Listener:**

| Tool | Adds |
|---|---|
| `add_aggregate_report_listener` | Aggregate Report listener |
| `add_summary_report_listener` | Summary Report listener |
| `add_view_results_tree_listener` | View Results Tree listener (full request/response capture for debugging) |
| `add_backend_listener` | Backend Listener (streams live metrics to InfluxDB/Graphite/etc.) |

**Editing** (mutate an already-built plan):

| Tool | Purpose |
|---|---|
| `remove_element` | Remove an element (and its subtree); rejects removing the root `TestPlan` node |
| `update_element` | Shallow-merge (or replace) a node's props; a prop value of `null` deletes that key. Validated against the node's type when known |
| `rename_element` | Rename an element's `testname` |
| `move_element` | Move an element (and its subtree) to a new parent, optionally at a specific index; rejects moving a node into its own subtree |
| `reorder_children` | Reorder a node's direct children (must pass an exact permutation of the current children) |
| `set_element_enabled` | Enable/disable an element without removing it |

**Inspection:**

| Tool | Purpose |
|---|---|
| `list_test_plans` | List every plan in the workspace |
| `get_test_plan` | Full element tree of a plan, including every node's `id` |
| `get_test_plan_xml` | Serialize a plan to its JMeter `.jmx` XML, without running JMeter |
| `import_test_plan` | Import an externally authored `.jmx` (e.g. exported from the JMeter GUI) as a new plan. Element types this server doesn't model are kept as opaque `UnknownElement` nodes instead of being dropped |

**Execution & reporting** (async — a run happens in the background):

| Tool | Purpose |
|---|---|
| `execute_test_plan` | Serialize to `.jmx` and run JMeter in non-GUI mode; returns `{ executionId }` immediately |
| `get_execution_status` | `running` / `completed` / `failed`, plus a tail of the JMeter log |
| `stop_execution` | Send `SIGTERM` to a running JMeter process |
| `get_execution_report` | Aggregated stats (per label + overall) parsed from the run's JTL output |

**Capacity search** (async — the search runs many real test executions in the background):

| Tool | Purpose |
|---|---|
| `find_breaking_point` | Find the concurrency level where a plan stops meeting its SLA, by running it over and over and driving one thread group's thread count; returns `{ searchId }` immediately |
| `get_breaking_point_status` | `running` / `completed` / `failed` / `stopped`, every round run so far with its load and metrics, and the final breaking point |
| `stop_breaking_point_search` | Abort a running search, keeping the bounds it had already established |

### Finding the breaking point

`find_breaking_point` answers the question JMeter itself has no feature for: *how
many concurrent users can this survive?* Instead of you adjusting the thread
count, re-running, reading the report and deciding again — roughly six to eight
tool calls per round, across the six to eight rounds a search needs — the server
runs the whole search itself and you poll one tool for the answer.

It works in two phases:

1. **Bracketing** — starts at `startThreads` and doubles the load (50 → 100 →
   200 …) while the SLA holds, until a round breaks it or the `maxThreads`
   ceiling is reached.
2. **Binary search** — bisects between the last healthy level and the first
   broken one until the two are within `toleranceThreads` of each other, or
   `maxIterations` rounds have run, whichever comes first.

```
find_breaking_point (planId, threadGroupNodeId, maxThreads: 400, p95Ms: 800, errorPct: 1)
                                → { searchId }
get_breaking_point_status (searchId)   ← poll
                                → { breakingPoint: 137, lastHealthy: 134,
                                    breakingPointRange: { healthyUpTo: 134, brokenAt: 137, exact: false },
                                    stopReason: "converged", progress: {...}, iterations: [...] }
```

`breakingPoint` is the lowest load that was actually *tested* and broke the SLA.
Because the search stops bisecting at `toleranceThreads`, the levels between
`lastHealthy` and `breakingPoint` were never run — read `breakingPointRange` for
the real precision (`exact: true` only when the two are adjacent), and treat the
edge as "between 134 and 137", not "137". Status also reports `progress` (rounds
done, and the in-flight round's elapsed time and percent complete) and `files`
(where `meta.json` and the per-round executions live). There is no completion
push: poll `get_breaking_point_status`.

A round passes when **every** threshold set (`p95Ms`, `errorPct`) is met by the
overall `TOTAL` row; at least one threshold is required. Each round puts the
thread group into scheduler mode — a ramp-up proportional to the thread count
(`rampSecondsPerThread`) followed by a fixed `plateauDurationSeconds` at full
load — and **only the plateau samples count**, so the ramp-up doesn't drag the
numbers of a healthy round down. Loop counts are deliberately not used: they
would make rounds at different thread counts incomparable.

Because a round repeats the thread group's whole scenario until the plateau
ends, **the request mix per round follows from how the plan is built**: a login
that should happen once per user belongs under a Once Only Controller
(`add_once_only_controller`), or it will run on every repeat. Each round reports
its metrics overall *and* per label (`byLabel`, with each label's share of the
samples), so a skewed mix is visible right away — the SLA itself is judged on
the overall numbers.

The search temporarily overwrites `numThreads`, `rampTimeSeconds`,
`durationSeconds` and `loops` on the thread group you point it at, and restores
the original values when it ends — including when it fails or is stopped. The
status reports this as `propsRestored`, and the original values are kept in the
search's `meta.json` in case the process dies mid-search. The plan needs an
Aggregate Report, Summary Report, or View Results Tree listener, or there would
be no metrics to judge the SLA against.

| Parameter | Default | Purpose |
|---|---|---|
| `maxThreads` | *(required)* | Safety ceiling — the search never runs more threads than this |
| `p95Ms` / `errorPct` | *(at least one required)* | SLA thresholds a round is judged against |
| `startThreads` | `50` | Load for the first round |
| `toleranceThreads` | 2% of `maxThreads`, min 5 | How close the bounds must get before the search stops bisecting |
| `rampSecondsPerThread` | `0.1` | Ramp-up seconds per thread, so every round adds load at the same rate |
| `plateauDurationSeconds` | `60` | Seconds at full load; only these samples count toward the SLA |
| `cooldownSeconds` | `5` | Pause between rounds so the system under test recovers |
| `maxIterations` | `8` | Hard cap on rounds, so a slow search still ends |

A search that never breaks the SLA reports `breakingPoint: null` with
`stopReason: "ceiling-reached"` — raise `maxThreads` and go again. One where
even the first round breaks reports `lastHealthy: null`; lower `startThreads`.

While an execution runs, `get_execution_status` also returns a `progress` block
computed from the results file so far — samples, error rate, avg/p95 latency,
overall and recent throughput, and the same per label — so polling it is enough;
there is no need to tail `stdout.log`. Once the run ends, `get_execution_report`
gives the full aggregate.

### JSR223 scripts need a compatible Java

JSR223 elements default to Groovy, and the Groovy bundled with JMeter 5.6.x
(3.0.x) fails on very new Java releases: every script run throws, which ends
that virtual user's iteration and can look like load that never grows. Use Java
17 (LTS) for JMeter — set `JAVA_HOME` in the MCP server's environment. The
`add_jsr223_*` tools, `execute_test_plan` and `find_breaking_point` detect the
Java and Groovy versions in use and return a `warning` / `warnings` entry when
they don't match, and the tool descriptions steer clients toward the built-in
`${__UUID}`, `${__RandomString}` and `${__Random}` functions, which need no
script.

## Example workflow

```
create_test_plan            → { planId, rootNodeId }
add_thread_group             (parentId: rootNodeId)  → { nodeId: threadGroupId }
add_http_sampler              (parentId: threadGroupId) → { nodeId: samplerId }
add_response_assertion        (parentId: samplerId)
add_aggregate_report_listener (parentId: threadGroupId)
execute_test_plan             (planId) → { executionId }
get_execution_status           (executionId)   ← poll until "completed"
get_execution_report            (executionId) → aggregated latency/error stats
```

## Testing

212 automated tests, no framework beyond Node's built-in test runner:

```bash
npm test               # 196 tests: tree-mutation and XML-shape unit tests, XML -> tree parsing,
                        # serialize -> parse round-trips, the breaking-point search algorithm
                        # driven round-by-round against a simulated system, and every tool called
                        # over the real MCP protocol (stdio, the same way Claude Code/Desktop talk
                        # to it) - no JMeter install needed, fully hermetic
npm run test:integration  # 16 tests: real JMeter runs - the If Controller story above, While
                        # Controller loop counts, timer pacing, extractors, assertions, and a full
                        # find_breaking_point search converging on the real concurrency limit of a
                        # live service (needs JMETER_HOME)
npm run test:all
```

`npm test` spawns the actual built server (`dist/index.js`) via `StdioClientTransport` and drives it exactly as a real client would — not just calling internal functions — so a broken tool schema or a malformed response shows up as a real protocol error, not a passing unit test.

Both suites run on every push and pull request via [GitHub Actions](.github/workflows/ci.yml) — the integration job installs a real JMeter binary on the runner, so it's exercising the same code path as a local run, not a mock.

## Prerequisites

- Node.js 18+
- JMeter installed locally, with the `JMETER_HOME` environment variable
  pointing at the installation directory (the one containing `bin/jmeter`).
  On macOS via Homebrew, `brew install jmeter` puts it at
  `/opt/homebrew/opt/jmeter/libexec`.

## Adding this server to Claude Code

### Via npx (recommended — published on npm)

No cloning or building required; `npx` fetches and runs the published
version on the fly:

```bash
claude mcp add jmeter \
  -e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
  -- npx -y jmeter-mcp-server
```

Adjust the `JMETER_HOME` path to wherever JMeter is installed on your
machine. Optionally set `JMETER_MCP_WORKSPACE` too (see below) if you want
plans and executions stored somewhere other than the default.

The default scope is `local` (this project directory only). To make it
available across every project, add `-s user`:

```bash
claude mcp add jmeter -s user \
  -e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
  -- npx -y jmeter-mcp-server
```

Confirm it registered and is responding:

```bash
claude mcp list
```

### From a local clone (development)

If you're working on this repository's code instead of using the published
package, point at the built `dist/index.js` directly:

```bash
npm install
npm run build
claude mcp add jmeter \
  -e JMETER_HOME=/opt/homebrew/opt/jmeter/libexec \
  -- node /absolute/path/to/jmeter-mcp-server/dist/index.js
```

### Claude Desktop

Add this to `~/Library/Application Support/Claude/claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "jmeter": {
      "command": "npx",
      "args": ["-y", "jmeter-mcp-server"],
      "env": {
        "JMETER_HOME": "/opt/homebrew/opt/jmeter/libexec"
      }
    }
  }
}
```

Note: unlike a terminal-launched app, Claude Desktop does **not** inherit
environment variables exported in your shell profile (`.zshrc`, etc.) — only
true system-wide ones. Always set `JMETER_HOME` explicitly in the `env`
block above rather than relying on it already being "set on your machine".

## Environment variables

| Variable | Required | Purpose |
|---|---|---|
| `JMETER_HOME` | Yes | JMeter installation directory (must contain `bin/jmeter`) |
| `JMETER_MCP_WORKSPACE` | No | Where plans and executions are stored. Defaults to `./jmeter-workspace` relative to wherever the server process starts, so when the server runs inside a project repo, point this outside the repo (or gitignore `jmeter-workspace/`) to keep run artifacts out of `git status` |

## Workspace layout

```
<workspace>/
  plans/<planId>/plan.json           # JSON tree — source of truth for a plan
  executions/<executionId>/
    generated.jmx                    # serialized at execute_test_plan time
    aggregate-report.jtl             # output of the Aggregate Report listener, if present
    summary-report.jtl               # output of the Summary Report listener, if present
    jmeter.log
    meta.json                        # execution status, pid, timestamps, exit code
  capacity-searches/<searchId>/
    meta.json                        # find_breaking_point search: config, SLA, every round run,
                                     # current bounds, and the thread group's original settings
```

A `find_breaking_point` search doesn't get its own execution directory — each of
its rounds is a normal execution under `executions/`, and the search's
`meta.json` records the `executionId` of every round, so any individual round can
still be inspected with `get_execution_report`.

## Editing and importing plans

Beyond the `add_*` authoring tools, a plan can be mutated after the fact
(`remove_element`, `update_element`, `rename_element`, `move_element`,
`reorder_children`, `set_element_enabled`) and an externally authored `.jmx`
(e.g. exported from the JMeter GUI) can be brought in with `import_test_plan`.
`import_test_plan` understands the most common element types (thread groups,
HTTP samplers, assertions, extractors, controllers, config elements, the
three report listeners, etc.); anything it doesn't recognize is kept as an
opaque `UnknownElement` node whose original XML is preserved and re-emitted
as-is by `get_test_plan_xml`/`execute_test_plan`, instead of being dropped -
`import_test_plan`'s response reports `unknownElementCount`/
`unknownElementTypes` so you know what wasn't fully understood. Coverage can
be extended incrementally in `src/jmx/parser.ts`.

## v1 scope

Not yet supported (candidates for a future release): generating the HTML
dashboard report (`-e -o`), parent-type validation on `add_*`/`move_element`/
`import_test_plan` (nothing stops attaching an element under a semantically
wrong parent), distributed execution.

Note on `add_csv_data_set`: the `filename` must be an absolute path.
`execute_test_plan` runs JMeter from a fresh per-execution directory, so a
relative path (which JMeter's GUI would resolve against the `.jmx` file's
own location) won't resolve there. An absolute path baked into a plan is
also machine-specific — it won't travel if you share `plan.json` with
someone on a different machine. This is only checked at creation time:
`add_csv_data_set` rejects a relative or nonexistent path up front, but
later changing a `CSVDataSet`'s `filename` via `update_element`, or
importing a `.jmx` that already has a relative one via `import_test_plan`,
is not checked - it will only surface as a failure at `execute_test_plan`
time.

Note on `add_jdbc_request`/`add_jdbc_connection_configuration`,
`add_ftp_request`, and `add_backend_listener`: these generate correct,
JMeter-loadable XML, but exercising them for real needs infrastructure this
project doesn't provide (a database, an FTP server, an InfluxDB/Graphite
instance) — they were verified structurally, not against a real backend.

Note on `add_view_results_tree_listener`'s `captureFullData` option: it has
no effect right now. `execute_test_plan` always runs JMeter with
`-Jjmeter.save.saveservice.output_format=csv`, and JMeter's CSV writer never
emits response body/header columns no matter what the `SampleSaveConfiguration`
flags say — only its XML output format can carry full response bodies. The
option is wired up correctly in the generated `.jmx` (verified: the flags
really do flip in the XML) for the day this server supports XML-format runs,
but until then it's a no-op — confirmed by running a real capture and
checking the resulting JTL has no `responseData`/`samplerData`/
`requestHeaders`/`responseHeaders` columns regardless of the setting.

Note on `add_tcp_sampler`: `server`/`port`/`request` are live-verified. The
numeric fields (`connectTimeoutMs`, `timeoutMs`) are rendered as
`stringProp` following this project's general convention for sampler
numeric fields, but that specific choice for `TCPSampler` wasn't confirmed
against a real JMeter-GUI-saved example (none was available to check
against) — flagging in case a real save turns out to expect `intProp`.

## Roadmap

Ideas being explored for future releases — none of these are implemented yet:

| Proposed tool | What it does | Why it's worth it |
| --- | --- | --- |
| `detect_bottleneck_class` | Fits Little's Law / the Universal Scalability Law to collected concurrency vs. throughput vs. latency data, and classifies the bottleneck as contention, coherency, or saturation | JMeter only outputs raw numbers; re-deriving a queueing-theory curve fit through prose reasoning would mean reimplementing nonlinear regression by hand for every question |
| `detect_soak_drift` | Runs linear regression over the latency/error time series of a long-duration soak test to separate normal noise from a real trend (the classic memory-leak signal) | JMeter's graph shows the curve, but doesn't say whether it's a statistically real degradation or just noise |
| `compare_execution_reports` | Statistical diff across N executions (not just two), with a significance test for whether a p95 shift is real or noise | A ready-made performance regression gate for CI, instead of someone eyeballing two JSON reports and guessing |
| `isolate_warmup_window` | Automatically detects where warm-up (JIT, connection pools, cold caches) ends, and recomputes metrics only over the steady-state window | Today's overall average is polluted by the first few seconds of a run; JMeter doesn't separate this on its own |
| `classify_error_flakiness` | Re-runs failed samples with backoff and separates "real system error" from "one-off flake" (network timeout, etc.) | Produces a trustworthy error rate for a CI gate, instead of an `errorPct` that mixes both kinds together |
| `orchestrate_distributed_run` | Runs the same test plan across multiple JMeter injector nodes (master-slave, via `-R` or independent engines) and merges every node's `.jtl` into a single aggregated report | JMeter supports distributed mode, but wiring up remote engines, RMI ports, matching JMeter versions, and merging results across machines is entirely manual today — nobody sets this up for a quick test |

## How this compares

Checked against the other public JMeter MCP servers found on GitHub as of September 3, 2026 —
open-source repositories with at least a README description (undocumented forks/clones excluded).
Columns reflect what each project's own README documents, not independent verification of its
internals.

| Server | Approach | `.jmx` round-trip | Edit by id | Async run | Tests | Real JMeter |
| --- | --- | --- | --- | --- | --- | --- |
| **jmeter-mcp-server** (this project) | JSON tree, 34 element types, authored and edited by node id | ✓ | ✓ | ✓ | 212, incl. real runs | ✓ |
| [QAInsights/jmeter-mcp-server](https://github.com/QAInsights/jmeter-mcp-server) | Runs an existing `.jmx` and analyzes results — doesn't author plans | — | ✗ | ✗ | ✗ | ✓ |
| [aravindksk7/Jmeter-MCP](https://github.com/aravindksk7/Jmeter-MCP) | Generates a `.jmx` from parameters; no re-import of existing plans | ✗ | ✗ | ✗ | ✗ | ✓ |
| [chandanvars/jmeter-mcp-server](https://github.com/chandanvars/jmeter-mcp-server) | Generates a whole plan from one JSON payload, runs it via Docker | ✗ | ✗ | ✗ | ✗ | ✓ |
| [perfsage/perfsage-jmeter-mcp](https://github.com/perfsage/perfsage-jmeter-mcp) | Heals the local JMeter runtime, imports HAR/OpenAPI traffic, discovers capacity | — | ✗ | ✗ | ✗ | ✓ |
| [shruthi-r18/ClaudeCode_MCP_QA_Automation_Performance](https://github.com/shruthi-r18/ClaudeCode_MCP_QA_Automation_Performance) | Demo project for a tutorial video, 11 fixed tools | ✗ | ✗ | ✗ | ✗ | ✓ |
| [KenLin-7/jmeter-mcp-server](https://github.com/KenLin-7/jmeter-mcp-server) | Reimplements HTTP load testing in Node — no Apache JMeter underneath | ✗ | ✗ | ✗ | ✗ | ✗ |
| [MUYU0615/jmeter_mcp_server](https://github.com/MUYU0615/jmeter_mcp_server) | Generates a `.jmx` with basic parameters (threads, ramp-up, duration) | ✗ | ✗ | ✗ | ✗ | ✓ |
| [vjgit-369/JmeterDemoUsingMCP](https://github.com/vjgit-369/JmeterDemoUsingMCP) | Demo script with parameters hardcoded in source, not a general-purpose server | ✗ | ✗ | ✗ | ✗ | ✓ |

Every capability in that table shows up somewhere across the other eight projects, individually:
real JMeter execution, `.jmx` generation, traffic import. This is the only one that combines full
`.jmx` round-trip (import *and* export), per-id incremental editing, non-blocking async execution,
and a documented automated test suite in one place.

## License

MIT

TDQS

A3.6/5.0

Scored across 56 tools

Disambiguation5/5

Each tool targets a distinct JMeter element or operation: add_* tools for different element types (thread groups, samplers, controllers, timers, assertions, extractors, listeners, config elements), and separate tools for lifecycle operations (create, update, move, remove, execute, report). Even similar tools like add_json_extractor vs add_regex_extractor are clearly differentiated by response format. There is no ambiguity in selecting between them.

Naming Consistency5/5

All tool names follow a strict verb_noun snake_case pattern: add_* for creation, get_* for retrieval, execute/stop for control, etc. The verb is always first and lowercase, and nouns are descriptive (e.g., add_http_sampler, set_element_enabled, get_execution_report). No camelCase or mixed conventions are present.

Tool Count2/5

With 56 tools, the surface is significantly larger than typical MCP servers (calibration marks 25+ as excessive). While JMeter's complexity justifies a broad set, the count still feels heavy and might overwhelm agents. Some tools could be consolidated (e.g., different extractors could be parameterized), but as-is, the tool count is beyond the 'reasonable' range.

Completeness4/5

The tool set covers the full JMeter workflow: plan creation/import, element manipulation (add, remove, update, move, enable), execution control (execute, stop, status), and result reporting (aggregate, summary, view results tree). It includes a wide variety of samplers, controllers, timers, assertions, and extractors. The main gap is the lack of a tool to delete a test plan entirely (root node cannot be removed), leaving plans to accumulate. Minor gap, but otherwise comprehensive.

Maintenance

ActivityMaintained
ResponsivenessResponsive