Skip to main content
Glama
juliodelimas

jmeter-mcp-server

by juliodelimas

find_breaking_point

Find the concurrency level at which a JMeter test plan breaches its SLA by doubling load until failure, then binary-searching between healthy and broken levels. Returns a searchId to poll for results.

Instructions

Find the concurrency level where a test plan stops meeting its SLA. Runs the plan over and over on its own, driving one thread group's thread count: first doubling the load until the SLA breaks, then binary-searching between the last healthy level and the first broken one. Returns immediately with a searchId; poll get_breaking_point_status for progress and the final breaking point. The thread group is temporarily switched to ramp-up + fixed-plateau (scheduler) mode for the search and its original settings are restored when the search ends. Requires an Aggregate Report, Summary Report, or View Results Tree listener in the plan. Rounds always run the thread group in scheduler mode with an infinite loop count, so every thread repeats its whole scenario until the plateau ends: a request meant to run once per user (e.g. a login) must sit under a Once Only Controller, and a nested Loop Controller runs its count on every repeat. Each round reports its metrics both overall and per label (byLabel, with each label's share of the samples) - check that the mix matches what the plan intends. The SLA is judged on the overall numbers, not per label.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
p95MsNoFail a round when overall p95 latency exceeds this many milliseconds.
planIdYes
errorPctNoFail a round when the overall error rate exceeds this percentage.
maxThreadsYesSafety ceiling - the search never runs more threads than this.
startThreadsNoThread count for the first round; doubles from there while the SLA holds (default 50).
maxIterationsNoHard cap on rounds, so a slow search still ends (default 8).
cooldownSecondsNoPause between rounds so the system under test recovers (default 5).
toleranceThreadsNoStop bisecting once the healthy and broken bounds are this close (default: 2% of maxThreads, minimum 5).
threadGroupNodeIdYesNode id of the thread group whose thread count the search will drive.
rampSecondsPerThreadNoRamp-up seconds per thread, so every round adds load at the same rate (default 0.1).
plateauDurationSecondsNoSeconds to hold full load after ramp-up; only these samples count toward the SLA (default 60).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.4.5

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it delivers extensively. It reveals that the plan runs repeatedly, the thread group is temporarily switched to scheduler mode and restored afterward, rounds use an infinite loop count, and SLA judgment is based on overall metrics rather than per-label. These are non-obvious side effects and semantics that an agent must know before invoking the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place: purpose, algorithm, return behavior, mode-switching side effect, listener prerequisite, loop-count caveat, and metric semantics. It is front-loaded with the core purpose and progressively adds necessary operational detail without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with no output schema, the description is remarkably complete. It explains the asynchronous return pattern, the required listener, the thread-group mutation and restoration, the per-label vs overall metric distinction, and the SLA evaluation basis. An agent has enough information to invoke the tool correctly and interpret its behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 91%, so the schema already documents most parameters well. The description adds useful context about how parameters relate to the search algorithm (e.g., doubling, binary search, overall SLA), but it does not substantially elaborate individual parameter meanings beyond what the schema provides. This matches the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Find the concurrency level where a test plan stops meeting its SLA.' It then details the exact algorithm (doubling then binary search), which distinguishes it clearly from sibling execution and status tools. The purpose is unambiguous and not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational guidance: it returns a searchId to poll via get_breaking_point_status, and it states the prerequisite that an Aggregate Report, Summary Report, or View Results Tree listener must exist. It does not explicitly contrast with execute_test_plan or normal test runs, but the specialized search behavior and polling instructions make the intended usage evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.