MCP Failure Lab
Use this MCP server to test MCP client failure and recovery behavior with deterministic fault-injection tools over stdio or HTTP.
ping: check server responsiveness; no inputs.protocol_ping_liveness: send a protocol ping during an in-flight call and enforce bounded liveness; configurepingAfterMs,livenessTimeoutMs, optionalcloseOnFailure, and optionalcompletionDelayMs.delay: delay a successful response bydelayMs(0–30000 ms) to test client timeouts.hang: never respond unless the MCP request is cancelled.disconnect: close the MCP transport before the request receives a response.malformed_message: return one deterministic invalid JSON-RPC response usingvariant:missing-jsonrpc,invalid-jsonrpc-version, orresult-with-error.duplicate_response: send the same JSON-RPC response twice for one request.response_after_cancellation: send a late response for a cancelled stdio request.
Allows MCP Failure Lab scenarios to be run against GitHub's official remote MCP server, exercising GitHub tools through configurable failure and resilience tests.
Allows MCP Failure Lab scenarios to be run against GitLab through an OAuth-capable stdio MCP bridge, enabling fault-injection and resilience testing of GitLab MCP interactions.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Failure Lablist the planned fault injection scenarios"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Failure Lab
Reproduce MCP timeouts, cancellation races, transport loss, and invalid responses with repeatable tests and CI reports.
Documentation · Compatibility · Project page
Example output (duration varies):
$ npx mcp-failure-lab demo
MCP Failure Lab — Demo
Running a real 500ms delay scenario...
Scenario: Deterministic delay demo
Outcome: success
Duration: ~500 ms
Assertions: passedWhy this exists
Real MCP clients behave differently when things break. A timeout may leave a connection usable; an interrupted response may close it. A malformed reply may be rejected by one client and accepted by another.
Failure Lab makes those cases repeatable so you can check both the failed call and what happens next.
Related MCP server: mcp-chaos-rig
Real-world findings
Python SDK #3522: Python
mcp2.2.0 stayed closed after an interrupted HTTP response, rejected the next request, and raised anExceptionGroupduring cleanup.Rust SDK #1283:
rmcp3.4.0 accepted an invalid response containing bothresultanderrorover stdio and HTTP.Ruby SDK #589:
mcp1.6.1 accepted responses missingjsonrpcor declaringjsonrpc: "1.0"over stdio and HTTP in all three repeats. The Ruby report records all 60 executions; the results and upstream finding are included in the Observatory export.Java SDK 2.0.1 accepted
jsonrpc: "1.0"over both transports (#1156). Over stdio, responses missingjsonrpcor containing bothresultanderrorleft the nextpingtiming out (#1157). See the Java report.Duplicate-response comparison: all five tested SDKs completed the next
ping. TypeScript reported the duplicate through its error callback; the other harnesses surfaced no call-level duplicate error.
The SDK comparisons describe specific tested versions, not every release. See the versioned reports for reproduction steps and limitations.

Quick start
Requires Node.js 22.19.0 or newer and npm.
npx mcp-failure-lab demoThe demo needs no API key, external server, or global installation. To test your own MCP client, connect it to Failure Lab over stdio or local Streamable HTTP:
npx mcp-failure-lab serve
# Or:
npx mcp-failure-lab serve --transport httpThe stdio process waits for a client; it is not an interactive terminal command. Configure your
client to launch it, or follow the getting started guide.
Press Ctrl+C to stop a manually started server.
HTTP defaults to http://127.0.0.1:3000/mcp. It provides neither authentication nor TLS;
do not expose it to an untrusted network.
Test your own MCP server
External targets support HTTP and stdio. After the
repository setup,
install the official GitHub MCP server with its executable on PATH and set
GITHUB_PERSONAL_ACCESS_TOKEN in your environment. Then run:
npm run dev -- run examples/scenarios/github-get-me.json \
--target examples/targets/github-stdio.jsonThe example calls GitHub's read-only get_me tool. Its target configuration uses envFrom
to pass the token from your environment rather than storing it in JSON.
For your own scenario and
target configuration, use:
npx mcp-failure-lab run scenario.json --target target.jsonExternal runs check tool results, deadlines, and adapter setup and cleanup. They do not inject faults into another server; Failure Lab is not a proxy.
See the external targets guide for prerequisites, credential handling, and configuration.
What Failure Lab can break
The built-in server exposes these tools for testing client behavior:
Tool | Behavior |
| Returns a deterministic health response |
| Sends a bounded protocol ping during its in-flight call |
| Waits for a bounded duration before returning |
| Remains pending until the client cancels |
| Interrupts the active transport while a request is in flight |
| Violates one selected JSON-RPC response rule exactly once |
| Sends the same JSON-RPC response twice for one request |
| Sends one late response for a cancelled stdio request |
| Invalidates the caller's legacy HTTP session |
Protocol ping is distinct from the ping tool. Built-in scenarios default to MCP 2026-07-28;
protocol liveness requires scenario protocolVersion: "2025-11-25".
Both public transports accept legacy clients.
See the fault tools reference for timing bounds, activation, cancellation, cleanup, and transport limits.
How recovery testing works
Trigger a fault, record its outcome, then make a second request on the same connection. A detected fault does not prove recovery; the follow-up must succeed too.
Scenario observe calls run after the primary call, including errors and timeouts.
For example, duplicate-response.json
calls duplicate_response, then uses ping to check that the connection remains usable:
# From a repository checkout
npm run dev -- run examples/scenarios/duplicate-response.jsonThis checks post-fault behavior, not an automatic recovery policy. A stdio disconnect terminates the server process and requires a new process and connection.
See Scenarios for outcome, duration, result, and observer assertions.
SDK / transport evidence
MCP Failure Observatory collects SDK and transport evidence. The repository reports preserve tested versions, methods, and limitations.
Results vary by SDK, version, transport, and protocol. A passing run is evidence for that combination, not a guarantee for every client.
CI usage
Save a scenario file with explicit expectations and timeouts, then produce a JUnit report:
npx mcp-failure-lab run scenario.json --report junit > junit.xmlPin the Failure Lab package version in CI so upgrades do not change the test environment unexpectedly.
Exit codes: 0 means expectations passed, 1 means the scenario could not be loaded or executed,
and 2 means an assertion failed. An expected timeout can pass; an unexpected success can fail.
See Reporting for JSON, JUnit, and lifecycle diagnostics.
Documentation
Detailed guides live at mcplab.dev: Getting started · CLI · Architecture · Troubleshooting.
Contributing
See CONTRIBUTING.md for setup, tests, and the contribution workflow. Planned work is tracked in GitHub Issues.
If Failure Lab helps you test an MCP integration, consider starring the repository.
License
Available Tools
8 toolsdelayA
Delay a successful MCP response to test client timeout behavior.
| Name | Required | Description | Default |
|---|---|---|---|
| delayMs | Yes | Response delay in milliseconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose the key behavioral trait: the response is delayed but still successful, which meaningfully separates it from 'hang' (no response) and 'disconnect' (dropped connection). It adds no information about blocking behavior, the 30s ceiling, or whether the call is cancellable mid-delay.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the action and ends with the purpose. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter test utility with no output schema, the description covers the essentials: what is delayed and why. It would be fully complete if it noted that the call blocks the connection for the given duration and that the schema-bound maximum is 30 seconds.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter 'delayMs' is fully documented with units, minimum, and maximum in the schema. The description adds no syntax or semantics beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delay a successful MCP response') plus the intent ('test client timeout behavior'), so the agent knows exactly what the tool does. It does not distinguish itself from siblings like 'hang' or 'disconnect', which is close in function, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its use case ('test client timeout behavior'), which is enough to infer that this is for fault-injection testing rather than production calls. However, it never says when to prefer this over 'hang', 'disconnect', or 'protocol_ping_liveness', leaving the sibling selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
disconnectA
Close the MCP transport before this request can receive a response.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the key behavioral trait: the transport is torn down before a response can be delivered, so the caller will not get a normal reply. It stops short of saying whether the closure is permanent, whether reconnection is possible, or what error surfaces to the caller.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action verb first and only the timing detail that matters. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with no annotations and no output schema, the description conveys the one thing an agent must know: no response will arrive. Minor gaps remain around session persistence and recoverability, but nothing essential to invoking it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty, so there is nothing for the description to disambiguate. Baseline 4 applies for a 0-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (close) and resource (the MCP transport), which unambiguously separates it from siblings like ping, delay, hang, and malformed_message that manipulate the session in other ways. An agent can tell what this does without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through the phrase 'before this request can receive a response', which signals it is a failure-injection/test tool, but it never says when to pick disconnect over hang, delay, or the other siblings. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_responseB
Return the same JSON-RPC response twice for this request.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the core behavior (duplicating the response) but omits operational details such as whether duplication is one-time or persistent, and whether it affects only this request or future ones.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. Appropriate for a zero-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a protocol-testing tool with no annotations and no output schema, the description is minimally viable but leaves important contextual gaps, such as when to choose this tool over siblings and whether the duplication has side effects on the connection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is empty and there is nothing for the description to clarify. Baseline for zero-parameter tools is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and behavior: return the same JSON-RPC response twice for the request. It is clear what the tool does, but it does not differentiate itself from siblings like response_after_cancellation or malformed_message.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The agent must infer its role in the testing suite.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hangB
Never respond unless the MCP request is cancelled.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the core behavior (no response until cancellation) but omits what happens after cancellation (e.g., error, silent stop) and any side effects or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the key behavior, and no wasted words. Appropriately sized for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-param testing tool, the core behavior is stated, but missing usage context and post-cancellation behavior leaves the agent without full operational clarity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so baseline is 4. The empty schema requires no additional parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific behavior (never respond) and a condition (unless cancelled), which distinguishes it from siblings like delay or disconnect. However, it does not explicitly name alternatives or scope, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus siblings such as delay, disconnect, or malformed_message. The agent must infer its role in testing timeouts or cancellation scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
malformed_messageC
Return one deterministic response that violates a selected JSON-RPC rule.
| Name | Required | Description | Default |
|---|---|---|---|
| variant | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. "Deterministic" is a useful signal, but it never states whether the malformed frame corrupts or terminates the session, whether subsequent calls are affected, or how the client is expected to observe the violation. For a fault-injection tool, those side-effect semantics are the key information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though the brevity is partly under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a fault-injection tool with no annotations, no output schema, and an undocumented enum, the definition is too thin: it omits what the response looks like, the side effects on session state, and how each variant maps to a violated rule. An agent can invoke it but cannot predict the consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there is one required enum parameter, so the description should have compensated. It does not list or explain the variant values, though "selected JSON-RPC rule" gestures at the enum and the values (missing-jsonrpc, invalid-jsonrpc-version, result-with-error) are largely self-describing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete verb and output ("Return one deterministic response") and a concrete domain constraint ("violates a selected JSON-RPC rule"), which is enough to distinguish it from fault-injection siblings like duplicate_response or disconnect. It does not explicitly name those siblings, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites, and no routing to alternatives. An agent must infer that this belongs in protocol-conformance testing and that siblings cover the other fault classes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pingB
Check whether the MCP server is responsive.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden, but it does convey that the operation is a non-destructive responsiveness probe. It does not say what a response looks like, whether it blocks/times out, or whether a failure is surfaced as an error versus a negative result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no filler. It is appropriately sized for a trivial probe, though it leaves room to add a clause distinguishing it from the similarly named sibling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-param, no-output-schema liveness probe, the description covers the essential purpose but omits the result semantics (what 'responsive' returns) and the sibling disambiguation needed in this crowded sibling set. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate — the baseline of 4 applies. The description correctly implies a no-argument invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Check') and the target resource ('the MCP server') plus the intended signal (responsiveness), so the core purpose is unambiguous. However, it does not distinguish itself from the closely-named sibling 'protocol_ping_liveness', which an agent could easily confuse with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given, and no conditions for choosing this over 'protocol_ping_liveness' or any of the fault-injection siblings (delay, hang, disconnect) are provided. The agent must infer that this is the plain liveness check purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
protocol_ping_livenessC
Send an MCP protocol ping during this in-flight call and apply a bounded liveness policy.
| Name | Required | Description | Default |
|---|---|---|---|
| pingAfterMs | Yes | ||
| closeOnFailure | No | ||
| completionDelayMs | No | ||
| livenessTimeoutMs | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It mentions a 'bounded liveness policy' but never defines the bounds, the failure behavior, the effect of closeOnFailure, the role of completionDelayMs, or any side effects, leaving the tool's runtime behavior largely opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded sentence with no filler, which is structurally clean. For a four-parameter tool with no annotations or output schema, however, this level of brevity reads as under-specification rather than appropriate conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the four parameters, zero schema description coverage, no annotations, and no output schema, the one-sentence description is far too thin. It omits parameter semantics, timeout behavior, failure handling, and sibling differentiation, so an agent cannot safely invoke the tool from the description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for four parameters, and the description does not mention pingAfterMs, livenessTimeoutMs, closeOnFailure, or completionDelayMs at all. It therefore adds no meaning beyond the bare schema and fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The sentence names a specific verb and resource ('Send an MCP protocol ping') and adds a scope ('during this in-flight call'), so it is not pure tautology. However, 'apply a bounded liveness policy' is vague, and the description never distinguishes this tool from the sibling 'ping' or explains what the liveness policy actually does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use, when-not-to-use, or alternative guidance is given. The phrase 'during this in-flight call' hints at a use context, but with siblings like ping, hang, and disconnect available, the agent receives no routing information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
response_after_cancellationA
Send a response after this call is cancelled (stdio only).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the key behavior of sending a response after cancellation and restricts it to stdio, which are non-obvious traits, but it does not detail what response is sent or error behavior outside stdio.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words, covering action, timing, and transport constraint efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, zero-parameter test tool with no output schema or annotations, the description gives enough context to understand what it does and its transport limitation. It could optionally clarify the nature of the response or side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to document. This meets the baseline of 4 for a no-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (send a response) and timing condition (after cancellation), and adds a transport restriction (stdio only). It is clear, though it does not explicitly contrast itself with sibling tools like duplicate_response or disconnect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The stdio-only constraint implies usage context, but the description does not state when to use this tool versus alternatives or provide explicit preconditions. Usage must be inferred from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.12.0- First observed
delay - First observed
disconnect - First observed
duplicate_response - First observed
hang - First observed
malformed_message - First observed
ping - First observed
protocol_ping_liveness - First observed
response_after_cancellation
TDQS
Scored across 8 tools
Most tools target distinct failure modes (delay, hang, disconnect, malformed_message, duplicate_response, response_after_cancellation). Only ping and protocol_ping_liveness overlap somewhat as both concern liveness, but their descriptions differentiate normal responsiveness from protocol-level in-flight liveness.
All names are lowercase snake_case, but the set mixes bare verbs (ping, delay, hang, disconnect) with descriptive noun/adjective phrases (protocol_ping_liveness, malformed_message, duplicate_response, response_after_cancellation). This is readable but not a single predictable pattern.
Eight tools is well-scoped for a failure-injection lab. Each tool represents a distinct failure scenario or liveness check, so none feel redundant or out of place.
The set covers common failure modes: timeout/delay, hang, disconnect, malformed protocol messages, duplicate responses, and post-cancellation responses. Minor gaps remain, such as a valid JSON-RPC error response tool or out-of-order response injection, but core failure-lab needs are met.
Maintenance
Related MCP Connectors
MEOK MCP Test MCP — golden-file + schema-drift + tool-failure tests for any MCP server. Drop-in
MCP server for building and testing AI agents with multi-model experimentation and insights.
A Model Context Protocol server for Wix AI tools
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA comprehensive reference implementation demonstrating all features of the Model Context Protocol (MCP) specification, serving as documentation, learning resource, and testing tool for MCP implementations.1MIT
- AlicenseNot gradedqualityBmaintenanceA local MCP server that breaks on demand, allowing you to test your client against auth failures, disappearing tools, flaky responses, and token expiry from a web UI.54 npm10MIT
- AlicenseAqualityAmaintenanceConformance test harness for Model Context Protocol servers, validating JSON-RPC, transport, capability, schema, and more across multiple spec versions.1297 npmMIT
- AlicenseNot gradedqualityBmaintenanceA lightweight mock MCP server for local testing and resilience experiments, providing predictable tool responses with simulated latency and errors.1MIT