Mockzilla
OfficialMockzilla MCP lets agents manage and run Mockzilla API mock servers locally and in the cloud, without an account for local use.
Local setup: check if the Mockzilla CLI is installed, install it (download/go-install/go-run), and report bridge version/updates.
OpenAPI exploration: inspect, lint, and discover OpenAPI specs; search/read bundled Mockzilla docs.
Spec reshaping: simplify heavy specs and pack service directories into
.mockzarchives.Local mocking: serve specs/URLs/folders as ephemeral mock servers, call endpoints, mock single endpoints without a spec, list/stop/clear mocks, and inspect request history with diagnostics.
Replay & stability: record and replay responses per endpoint, pin generated responses, and list recordings.
GitHub publishing: check deployability, list user repos, publish mocks to a repo, wait for deployment, and unpublish/delete.
Hosted cloud (after login): log in/out, get org/access context, list sims, browse catalog products, deploy hosted mocks from catalog/spec/URL, and wait for live URLs.
Enables deploying a mock Stripe sandbox via mockzilla's catalog, allowing testing of Stripe integrations without a real account.
@mockzilla/mcp
MCP server for Mockzilla - an open-source API mock server for OpenAPI specifications. Let Claude Code, Claude Desktop, Cursor, or Gemini CLI install Mockzilla, inspect OpenAPI specs, and spin up realistic local mock APIs in seconds. No account required for local use.
Source: github.com/mockzilla/mockzilla-mcp
Use cases
Local API development - mock any OpenAPI spec without a real backend or sandbox account
CI/CD integration testing - zero external dependencies in your pipeline
PSP and payment API mocking - Stripe, PayPal, Adyen from your editor without test accounts
Crypto exchange API mocking - Binance, Bybit without registered accounts
Rate limit protection - develop against OpenAI, Twilio without burning quota
Agentic workflows - let Claude or Cursor spin up and manage mock servers automatically
Related MCP server: mcp-cli-catalog
Two planes
@mockzilla/mcp exposes two planes of tools to your MCP client.
Local plane (no account required)
Runs on your machine. No Mockzilla account needed.
From an agent you can:
Check whether the Mockzilla CLI is installed, and install it into a managed cache (no changes to system PATH).
Inspect and lint an OpenAPI spec, or scan a folder for specs.
Simplify a spec that is too heavy to mock, or pack services into a
.mockzarchive.Serve any OpenAPI spec locally as a portable mock server, and call its endpoints.
Mock a single HTTP endpoint without a spec.
List, stop, and clear locally managed mocks.
Hosted plane (log in once)
Ask your agent to log in, or ask for something hosted. The agent calls the login tool, your browser opens the Mockzilla login, and you pick an organization and read-only or read-and-write access. The bridge keeps the login on your machine and renews it by itself. Hosted tools are listed from the start; before you log in, the agent is told to log in first.
Agents can then:
List deployed sims.
Browse catalog products.
Deploy hosted mocks from a spec, URL, or catalog bundle.
Wait for a deploy and return the live URL.
Before logging in, only the local plane is exposed. Agents can still help users explore Mockzilla and run local mocks before they sign up.
Example prompts
You can use these directly from Claude Code, Claude Desktop, Cursor, or Gemini CLI once mockzilla is configured as an MCP server.
Local plane (no account)
"Is the mockzilla CLI installed on this machine?"
"Install Mockzilla for me."
"Spin up the Petstore OpenAPI spec locally so I can curl it."
"What endpoints does
https://example.com/openapi.yamlexpose?""Mock
POST /checkoutto return a 402 response.""List the mock endpoints you're managing."
"Stop the mock server you started."
Hosted plane (after logging in)
"Log me in to Mockzilla."
"List the sims I have deployed."
"Show me the catalog products."
"Deploy a Stripe mock named
stripe-testand give me the live URL.""Create a hosted mock from this OpenAPI URL on mockzilla.org."
Install
Claude Code
One-liner, no config file editing:
claude mcp add -s user mockzilla -- npx -y @mockzilla/mcp@latest-s userinstalls for your user account (available in every project).Drop
-s userto scope to the current project only.
Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"mockzilla": {
"command": "npx",
"args": ["-y", "@mockzilla/mcp@latest"]
}
}
}Restart Claude Desktop after editing.
Cursor
Easiest: Settings -> MCP Servers -> Add new MCP server and fill in:
Name:
mockzillaCommand:
npxArgs:
-y @mockzilla/mcp@latest
Or edit ~/.cursor/mcp.json directly:
{
"mcpServers": {
"mockzilla": {
"command": "npx",
"args": ["-y", "@mockzilla/mcp@latest"]
}
}
}Restart Cursor after editing.
Gemini CLI
One-liner, no manual JSON editing:
gemini mcp add -s user mockzilla npx -y @mockzilla/mcp@latest-s userwrites to~/.gemini/settings.json(available in every project).Drop
-s user(or use-s project) to scope to the current directory's.gemini/settings.json.
Or edit the settings file directly:
{
"mcpServers": {
"mockzilla": {
"command": "npx",
"args": ["-y", "@mockzilla/mcp@latest"]
}
}
}Restart the Gemini CLI after editing.
Why @latest?
Without @latest, npx caches the first resolved version and won't pick up new publishes. Pinning to @latest makes npx re-check the registry on every spawn, so a Claude / Cursor / Gemini restart is enough to upgrade. Trade-off: ~200 ms extra startup time.
Local tools
These tools are always available and never leave the user's machine.
Setup and status
check_cliResolve Mockzilla on this machine: systemPATH-> bridge cache ->go runinvocation. Returns install options if nothing matches.install_cliInstall Mockzilla into~/.cache/mockzilla-mcp/. Methods:download(prebuilt from GitHub releases, default),go-install,go-run. Never touches systemPATH.bridge_statusReport the bridge's version, check npm for newer publishes, and surface upgrade steps.
OpenAPI exploration and docs
infoSummarise an OpenAPI spec without serving it:{title, version, openapi_version, endpoint_count, paths}. Also reads.mockzpackages.lintFind schemas in a spec that no value can satisfy, so a broken spec is caught before serving it:{clean, defect_count, defects}.discover_specsScan a directory for OpenAPI specs and folders of static endpoint files. Returns asuggested_inputforserve_locally.mockzilla_docs_topicsList the Mockzilla docs, by category, with each topic's title and summary. The product docs from mockzilla.org and the engine docs both ship inside the package, so no network or login is needed.mockzilla_docs_readReturn the full markdown for one or more topics, or a whole category.mockzilla_docs_searchKeyword search across all docs; returns top sections with snippets.
Reshaping specs
simplifyDrop or reduce union types, stripx-*extensions, and optionally cap optional properties per schema. Writes the simplified spec to disk and returns its path.packPack a directory of services into a.mockzarchive thatserve_locallycan serve, even from a URL.
Local mocking
serve_locallyStart a portable mock server on a free port. Accepts a spec file, directory, or publichttpsURL. Returns{url, port, pid, services}. For a single spec,latency,errors,mountandcontextadd delay, inject error responses, change the mount path, or set replacement values.stop_locallyStop a server started byserve_locally.call_endpointMake an HTTP request and return{status, headers, body}, to show a mock's response. Localhost only unlessallow_remoteis set.mock_endpointQuickly mock a single HTTP endpoint without an OpenAPI spec. Writes a static response into the managed mocks dir and (re)starts the shared server. Takesstatusandheadersfor a failure or a redirect with a real body (404 with an error payload, 201 with aLocation); omitresponsefor a body-less 204 or 304. Needs mockzilla 2.8.20 or newer.list_mock_endpointsList all endpoints currently mocked, plus the running server's URL and the Mockzilla UI URL.clear_mock_endpointsWipe all mocks and stop the managed server.request_historyList the requests the running server answered: method, URL, status, content type, latency, and whether the response came from an upstream. Passidwithservicefor one request's full headers and body. Reads the server's own history API, so no account is needed.diagnose_requestsExplain what is wrong with the recorded traffic and where its data came from: a breakdown by origin, status and content type, latency p50/p95/max, and findings such as an upstream that failed and silently fell back to a generated mock, a 404 from a wrong mount prefix, or a JSON body under a non-JSON content type.setup_replayConfigure replay for a service: record a response once, then serve it back for every matching request. Returnsrecording_scopesaying what each endpoint is keyed by, since with no match fields every call to an endpoint shares one recording. Writesconfig.ymlonly inside the bridge's own mocks dir; for a folder from your own project it returns the YAML and where it goes.list_replaysList a service's recordings, with the request values each one is keyed by.check_github_deployableSay whether a repository or folder would deploy a mock, and what is missing if not. Knows both kinds: portable service folders and a codegen Go server. Read-only.list_github_reposThe user's repositories, flagging which already publish mocks, so the agent can ask which one to use.publish_to_githubPush mocks to one of the user's own repositories and let the Mockzilla action deploy them at a host of their own,https://<label>.api.mockz.io. The label is the repository name, with-2,-3if it is taken. No Mockzilla account needed: the first push registers the repository. Merges into an existing services folder rather than replacing it, and never overwrites an existing workflow.wait_for_github_deployWait for the workflow run and return the live URL, as the action reported it.unpublish_from_githubTake the mocks down and free the simulation slot, optionally deleting the repository too.
Account
loginOpens the Mockzilla login in the browser, where the user picks an organization and read-only or read-and-write access. Returns right away with the login link. The agent calls a hosted tool again once the user approves. The login is saved under~/.config/mockzilla-mcp/, one per server URL, and renewed automatically.logoutRevokes the connection and deletes the saved login. Log out and in again to switch organization or access.
Hosted tools
@mockzilla/mcp lists the hosted tools from the start and forwards them to mockzilla.org's MCP endpoint once you log in. At the time of writing, the hosted surface includes:
get_contextlist_simslist_catalog_productsdeploy_mock_from_catalogdeploy_mock_from_specdeploy_mock_from_urlwait_for_deploy
Refer to the hosted server's docs or the MCP registry entry for the live tool list.
On a machine without a browser, such as CI or a remote server, set MOCKZILLA_TOKEN to an API key from the dashboard instead of logging in. The hosted tools are then available from the start.
Configuration
Env var | Default | Purpose |
| unset | API key ( |
|
| Override the hosted endpoint, e.g. |
| unset | Set to |
|
| OAuth client id |
| matches bridge version | Pin a specific Mockzilla CLI version for |
|
| Preferred port for the |
| unset | Read docs from another build of |
Files
The bridge keeps the CLI and mocks under ~/.cache/mockzilla-mcp/, and the login under ~/.config/mockzilla-mcp/:
~/.cache/mockzilla-mcp/
├── bin/mockzilla # downloaded or go-installed binary
├── config.json # { method, version, invocation? }
└── mocks/ # mock_endpoint persists static endpoints here
└── services/
└── <first path segment>/<rest of path>/<method>/
├── index.<ext> # the response body
└── meta.json # status and headers, written only when set
~/.config/mockzilla-mcp/
└── credentials.json # saved logins, one per server URL, readable only by yourm -rf ~/.cache/mockzilla-mcpresets the CLI and all mocked endpoints. The login stays.To wipe just the mocks:
rm -rf ~/.cache/mockzilla-mcp/mocks.To drop the login, ask the agent to log out. That also revokes it on mockzilla.org.
The system
PATHis never touched, so reset doesn't affect a separatebrewinstall of Mockzilla.
Updates
Recommended way to stay current:
Pin
@mockzilla/mcp@latestin your MCP client config sonpxre-checks the registry on every spawn.Restart Claude Desktop / Cursor / Gemini periodically. That's when the new tarball is fetched.
If something seems off, ask the agent: "Run
bridge_statusand tell me if@mockzilla/mcpis up to date."
If it's stale, run:
npx clear-npx-cache @mockzilla/mcpand restart your MCP client.
The Mockzilla CLI version is pinned by the bridge (via MOCKZILLA_VERSION in lib/install.js). Updating the bridge updates the pin; the next install_cli call brings the CLI itself up to date.
Development
See CLAUDE.md for project conventions and a walkthrough of adding a new tool.
Releasing
The bridge has two registries to keep in sync: npm (@mockzilla/mcp) and the MCP registry (server.json). Skipping the second one leaves discovery clients pinned to the previous tarball.
The usual way is to bump version in package.json, merge, and publish a GitHub release tagged v<version>. That covers both registries: see From GitHub Actions below.
To release by hand instead:
Bump
versioninpackage.json.Run:
make publish-allThis will:
Build
docs/andhosted-tools.jsonfrom the published docs bundle, plus the engine docs at the pinned CLI version. Maintainers only: it needsDOCS_BUNDLE_CMDinlocal.mk.Run the smoke tests: the stdio round-trip, login against a fake OAuth server, the docs tools,
mock_endpointagainst the real CLI, and the behavior checks that each tool reports only what the server really does (CLI-dependent ones skip when no CLI is installed).npm publishthe new tarball.Mirror the version into
server.json.Log
mcp-publisherin with the GitHub token from Keychain (see below).Run
mcp-publisher publishagainst the MCP registry.
Commit the
server.jsonbump.
If you only want one side:
make publishfor npm only.make publish-mcpfor the MCP registry only (server.jsonis always re-synced frompackage.jsonfirst).
mcp-publisher must be on PATH (brew install mcp-publisher or follow the installation docs).
MCP registry login
A registry login lasts 5 minutes, so make publish-mcp logs in before every publish. It reads a GitHub token from macOS Keychain, because the browser login (mcp-publisher login github without a token) can't publish under io.github.mockzilla.
One-time setup:
Create a classic GitHub token with
read:organdread:user.Store it in Keychain. The command prompts for the token:
security add-generic-password -a "$USER" -s mcp-publisher-github -w
From GitHub Actions
Publishing a GitHub release whose tag matches package.json runs .github/workflows/publish.yml, which builds the docs, runs the smoke tests and publishes to npm with provenance. It skips a version npm already has, so it is safe to run after make publish.
Two things have to be set up for it, both one-time:
npm Trusted Publishing. On npmjs.com, under the package's settings, add a trusted publisher: repository
mockzilla/mockzilla-mcp, workflowpublish.yml. No token is stored anywhere.AWS_ROLE_ARNrepository secret, set to thegithub_mcp_release_role_arnoutput of infra'slive/aws/iam-github-oidc. The docs bundle is only in S3 - nothing serves it, and the CDN in front of that bucket caches for a day - so the release assumes a role that can read that one key.
.github/workflows/publish-mcp.yml starts on the same release and covers the other registry. It skips versions the MCP registry already has, waits up to 10 minutes for the version to appear on npm, then publishes server.json using GitHub OIDC, so no token is needed. It can still be started by hand from the Actions tab.
The two run in parallel rather than in sequence. The wait loop is what keeps that safe, and it is needed anyway: npm accepts a tarball before it can be read back.
Related
mockzilla.org - hosted API simulation, per-PR mock URLs, GitHub Actions integration
mockzilla/mockzilla - the core open-source API mock server
Documentation - full usage guide for portable and codegen modes
License
Copyright © 2026-present
Licensed under the MIT License
Available Tools
35 toolsbridge_statusA
Report the bridge's own version and check whether a newer one is on npm. Returns {bridge_version, bridge_latest, update_available, upgrade_steps}. Call this when the user asks 'is mockzilla-mcp up to date?', or proactively if a tool starts failing in a way that could be a stale-bridge issue.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Returns specific fields and is read-only. No annotations, but the description explains output and usage. Could mention external call but not required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, efficient: first states purpose and output, second gives usage guidance. No redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a simple status tool: describes action, return fields, and when to use. No output schema but description covers it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters; schema coverage 100%. Description adds meaning by specifying return fields and usage, meeting baseline for 0 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it reports bridge version and checks for updates on npm. Names the return fields and describes the action with specific verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to call when user asks if updates exist or proactively when tools fail due to stale bridge. Provides clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
call_endpointA
Make an HTTP request to a URL and return {status, headers, body}. Use this to demonstrate a mock by hitting it after serve_locally (e.g. http://localhost:PORT/openapi/pet/findByStatus), to inspect the admin API (/.services returns the registered services, /healthz for liveness), or to verify a freshly-mocked endpoint works. Default scope is localhost only; pass allow_remote: true for arbitrary URLs (rare — the bridge isn't a general-purpose HTTP client).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| body | No | ||
| method | No | GET | |
| headers | No | ||
| allow_remote | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals the return shape ({status, headers, body}), default localhost-only scope, and the rarity of allow_remote. However, it could mention rate limits or permissions for making external requests, but the explicit security constraint (localhost default) compensates, making it mostly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured paragraph of 4 sentences. It front-loads the main action (HTTP request) and follows with specific use cases and constraints. Every sentence adds value without redundancy, making it concise and easily scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema or annotations, the description covers all necessary aspects: return format, typical use cases, security scope, and limitation (not general-purpose). It is sufficient for an agent to understand and correctly invoke the tool, addressing the complexity of making an HTTP request.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% property descriptions, so the description must compensate. It provides meaning for url (example localhost URLs), method (implicit via enum), and allow_remote (rare). Headers and body are not elaborated beyond the schema, but the overall context of making an HTTP request gives them implicit meaning. The description adds value but does not fully detail all parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's core function: making an HTTP request and returning status, headers, body. It distinguishes itself from sibling tools like mock_endpoint or peek_openapi by providing specific use cases (testing mocks, inspecting admin API, verifying endpoints). This satisfies the specific verb+resource requirement and differentiates from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly outlines when to use the tool (demonstrate a mock, inspect admin API, verify mock endpoint) and when not to (the bridge is not a general-purpose HTTP client, allow_remote is rare). This provides clear context and exclusions, guiding the agent to appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_cliA
Check whether the mockzilla CLI is available — either on the system PATH, in the bridge's own cache (~/.cache/mockzilla-mcp/), or via a go run invocation. Call FIRST when the user wants to try mockzilla locally. If nothing resolves, the response carries install_options; suggest install_cli to the user and ask them which method (download / go-install / go-run) they prefer.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses search paths and response content (install_options), though does not detail whether it executes the CLI or just checks files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise and structured, though slightly verbose with user guidance that could be secondary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers key aspects: inputs (none), actions, and follow-up steps. No output schema, but description implies response structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters; baseline 4 applies. Description adds no parameter info but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks CLI availability across three specific locations, and distinguishes its use case as the first step for local mockzilla attempts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call FIRST' and provides follow-up action: if unavailable, suggest install_cli with method options.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_github_deployableA
Say whether a GitHub repository (or a local folder) would deploy a mock, and what is missing if not. Use it when the user has a repo and asks whether it can serve mocks, or before adding the action to one. Detects which of the two kinds it is: a portable repo of service folders, or a codegen repo that builds a Go server. Returns problems (these stop it deploying), warnings (it deploys but something costs later, such as no teardown trigger, or specs large enough to exhaust a free simulation's 128MB), and workflow_to_add when there is no Mockzilla workflow at all. Read-only: it never changes the repository.
| Name | Required | Description | Default |
|---|---|---|---|
| dir | No | A local folder to inspect instead. | |
| repo | No | "owner/name" of a repo to inspect. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it succeeds: it explicitly states the tool is read-only and never changes the repository, describes the shape of the return value (problems, warnings, workflow_to_add), and gives concrete examples of warnings such as missing teardown triggers and 128MB simulation limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: purpose first, usage second, then return behavior. Every sentence adds useful information, with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given two self-documenting parameters and no output schema, the description adequately covers what the tool does, when to use it, what it returns, and its safety profile. An agent can decide to invoke it and interpret its results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both parameters. The description adds no new parameter-level meaning beyond referring to repos and local folders, matching the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it says whether a repository or folder would deploy a mock and what is missing. This clearly distinguishes it from sibling tools like deploy_mock_from_url, wait_for_github_deploy, or publish_to_github.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance is provided: when the user has a repo and asks whether it can serve mocks, or before adding the action to a repo. It does not name exclusions or alternatives, but the context is clear enough for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clear_mock_endpointsA
Wipe ALL mocks created via mock_endpoint and stop the managed server. Equivalent to rm -rf ~/.cache/mockzilla-mcp/mocks plus stop_locally. Use when the user wants to start fresh. Does not touch the mockzilla CLI binary or other bridge state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully bears the burden. It describes the destructive nature (wiping all mocks and stopping server), provides an equivalent command (`rm -rf ~/.cache/mockzilla-mcp/mocks` plus `stop_locally`), and explicitly states what it does NOT affect ('Does not touch the mockzilla CLI binary or other bridge state'), offering full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured paragraph. It front-loads the primary action, includes a functional equivalent, and adds clarifying statements. Every sentence serves a purpose, with no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description provides complete context: what the tool does, how it relates to siblings (`mock_endpoint`, `stop_locally`), its effects, and its limitations. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is 100%. The description adds meaning by explaining the effect, but since no parameters exist, it cannot add parameter-level detail. Baseline 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Wipe ALL mocks created via `mock_endpoint` and stop the managed server.' It uses a specific verb ('Wipe') and resource ('ALL mocks'), and distinguishes itself from sibling tools like `mock_endpoint` (which creates) and `stop_locally` (only stops server).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use when the user wants to start fresh,' providing a clear usage context. It does not explicitly state when not to use, but the context is sufficient to infer that this is for complete cleanup; alternatives like `stop_locally` exist for partial actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deploy_mock_from_catalogA
Create a HOSTED, SHAREABLE mock from a catalog spec (Stripe, Adyen, etc.) on mockzilla.org. The mock persists in the user's account and gets a stable URL anyone with the link can hit. Use this when the user wants something durable, team-visible, or reachable from outside their machine — NOT for ephemeral local exploration (use serve_locally for that). Pass catalog_spec_id from list_catalog_products; mount_path defaults to the spec slug. The returned sim is in deploying state -- follow up with wait_for_deploy to receive the live URL.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| slug | Yes | ||
| domain | No | ||
| mount_path | No | ||
| version_id | No | ||
| catalog_spec_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only include readOnlyHint:false, so the description carries the behavioral burden. It adds significant context: the mock persists, gets a stable URL, and the returned sim is in 'deploying' state requiring a follow-up call. It does not disclose error handling, authentication requirements, or potential naming conflicts, but the added state transition detail is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: purpose, usage guidance, parameter hints, and follow-up. Front-loaded with the core action and scoping, no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 0% schema coverage, no output schema, and minimal annotations, the description covers the essential workflow (deploy, then wait). It explains the async nature and the need for a follow-up. However, it doesn't explain all required parameters, which would be necessary for fully autonomous correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies `catalog_spec_id` (from list_catalog_products) and `mount_path` (defaults to spec slug) but leaves `name`, `slug`, `domain`, and `version_id` unexplained. Since `name` and `slug` are required, this is a meaningful gap, though not total.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific verb 'Create' and resource: a HOSTED, SHAREABLE mock from a catalog spec on mockzilla.org. It distinguishes from siblings by specifying catalog-based deployment and contrasting with serve_locally, and the mention of catalog_spec_id differentiates it from deploy_mock_from_spec and deploy_mock_from_url.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly provides when to use ('durable, team-visible, or reachable from outside their machine') and when NOT to use ('NOT for ephemeral local exploration') with a named alternative (`serve_locally`). Also gives prerequisite guidance to pass `catalog_spec_id` from `list_catalog_products` and instructs the follow-up with `wait_for_deploy`.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deploy_mock_from_specA
Create a HOSTED, SHAREABLE mock on mockzilla.org from an inline OpenAPI 3.0+ spec (YAML or JSON in the spec field, 4MB cap). The mock persists in the user's account and gets a stable URL anyone with the link can hit. Use this when the user pastes spec content AND wants a durable, team-visible result — NOT for ephemeral local exploration (use serve_locally for that). The returned sim is in deploying state — follow up with wait_for_deploy to receive the live URL.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| slug | Yes | ||
| spec | Yes | ||
| domain | No | ||
| filename | No | ||
| mount_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=false annotation, the description discloses that the mock persists in the user's account, receives a stable shareable URL, has a 4MB cap, and starts in the 'deploying' state requiring a follow-up call. It does not disclose side effects like slug collisions or overwrites, but the main behavioral traits are clearly surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler. The core action and key constraints are front-loaded, usage guidance follows, and the state/follow-up note closes the loop. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and sparse input schema descriptions, the description covers the key workflow: input format, persistence, state, and next step. It is still missing explanations for several parameters and edge-case behavior, but the agent has enough context to invoke the tool correctly and proceed to wait_for_deploy.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the `spec` parameter format (YAML/JSON, 4MB cap), but gives no guidance for the required `name` and `slug` parameters or the optional `domain`, `filename`, and `mount_path`. This leaves most parameters semantically underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Create'), a specific resource ('HOSTED, SHAREABLE mock on mockzilla.org'), and the input ('inline OpenAPI 3.0+ spec'). It clearly distinguishes itself from siblings like serve_locally, deploy_mock_from_url, and deploy_mock_from_catalog by emphasizing inline spec content and durable team-visible results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states exactly when to use this tool ('when the user pastes spec content AND wants a durable, team-visible result') and when not to ('NOT for ephemeral local exploration'), explicitly naming serve_locally as the alternative. It also tells the agent the follow-up step with wait_for_deploy, leaving no ambiguity about the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deploy_mock_from_urlA
Create a HOSTED, SHAREABLE mock on mockzilla.org from an OpenAPI 3.0+ spec at a public https URL. The server fetches the spec (SSRF-protected) and parses it. Use this when the user gives you a spec URL AND wants a durable, team-visible result — NOT for ephemeral local exploration (use serve_locally for that, it accepts the same URL form). The returned sim is in deploying state -- follow up with wait_for_deploy to receive the live URL.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| slug | Yes | ||
| domain | No | ||
| spec_url | Yes | ||
| mount_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false, so the description carries the behavioral burden. It discloses that the server fetches the spec (with SSRF protection), parses it, returns the sim in a `deploying` state, and requires a follow-up call. This gives the agent real operational context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place: core action, usage condition + alternative, and required follow-up. There is no repetition or filler despite the richness of the context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core invocation path is fully covered: source format, URL requirements, target platform, deploy state, and next action. The only gap is the lack of explanation for optional parameters (`domain`, `mount_path`) and not naming `deploy_mock_from_spec`/`deploy_mock_from_catalog` as related alternatives, but these are secondary to correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description only clarifies `spec_url` (public https URL). It does not explain `slug`, `domain`, or `mount_path`, and those optional parameters are otherwise undocumented in the schema. The required `name` and `slug` are somewhat self-evident, but an agent has to guess at the behavior of the optional fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action ('Create', a hosted mock), a target resource (mockzilla.org, durable/team-visible), and a precise input ('OpenAPI 3.0+ spec at a public https URL'). It immediately distinguishes this from local/exploratory use, so an agent can tell it apart from serve_locally without deeper inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (user has a spec URL AND wants a durable, team-visible result) and when not to (ephemeral local exploration), naming the alternative tool serve_locally. It also instructs the agent to follow up with wait_for_deploy, which is essential for completing the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diagnose_requestsA
Explain what is going wrong with the traffic the local server has served, and where each response's data came from. Returns a breakdown by source (generated / upstream / replay / cache), by status and by content type, latency p50/p95/max, and a list of concrete findings: upstream failures that silently fell back to a generated mock, 404s from a wrong mount prefix, a body that is JSON under a non-JSON content type, and unusually slow requests. Prefer this over request_history when the user asks why a response looks wrong, whether data is real or mocked, or why something is slow. Note that a replayed or cached response never reaches the history log, so a call the user made and cannot find here was most likely served from a recording or the cache rather than not happening. Read-only, localhost only.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many of the newest requests to examine. | |
| service | No | Limit to one service. Omit to look at every service. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so excellently: it declares the operation read-only, scopes it to localhost, describes the output categories, and explains the important behavioral caveat that replayed or cached responses do not appear in the history log. Nothing about the tool's behavior is hidden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet every sentence earns its place: purpose first, then output details, then usage guidance, then a caveat, then constraints. It front-loads the core function and provides actionable detail without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description thoroughly explains what the tool returns, when to use it versus a sibling, and an important behavioral limitation. The parameter schema covers the inputs, and the description covers the behavior and interpretation of results, so an agent has everything needed to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents `limit` and `service` adequately. The description does not add extra semantic nuance to the parameters, but it does not need to do so given the schema's coverage; it stays at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Explain what is going wrong') with a clear resource ('the traffic the local server has served') and then enumerates concrete outputs: breakdown by source, status, content type, latency percentiles, and concrete findings. It also distinguishes itself from the sibling `request_history` by mentioning the preference rule, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs the agent to prefer this tool over `request_history` when diagnosing why a response looks wrong, whether data is real or mocked, or why something is slow. The note about replayed/cached responses never reaching the history log also helps the agent interpret missing results and choose this tool appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discover_specsA
Scan a directory and report what mockzilla can do with it: top-level OpenAPI spec files (with title and endpoint count) plus any folders of static endpoint files mockzilla can serve. Returns a suggested_input the agent can hand directly to serve_locally. Use this when the user says 'I have a folder of specs/files, what's in it?' or 'mock this directory'. Scans one level only: on a tree of spec folders it names the subdirectories to recurse into. A big folder is summarised up to a cap, with truncated: true and the full spec_file_count.
| Name | Required | Description | Default |
|---|---|---|---|
| dir | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the scan is shallow ('Scans one level only'), that large folders are capped with `truncated: true`, and that the full `spec_file_count` is still returned. It does not explicitly state read-only behavior or error cases, but 'scan and report' strongly implies non-destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then adds only high-value details: the returned `suggested_input`, exact user triggers, the one-level depth limitation, and truncation behavior. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter discovery tool with no output schema, the description covers what is scanned, what is returned, how to use the output downstream with `serve_locally`, and edge behavior for large folders. Minor omissions such as error handling or path requirements do not significantly harm an agent's ability to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has a single `dir` parameter with 0% description coverage, and the tool description does not explain what `dir` should contain beyond the phrase 'Scan a directory'. There is no guidance on path format, whether the directory must exist, or whether absolute/relative paths are accepted, so the agent must infer semantics from the parameter name alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Scan'), a resource ('a directory'), and a concrete outcome ('report what mockzilla can do with it'). It also clarifies the return value (`suggested_input`) and distinguishes the tool from siblings like `serve_locally` by positioning it as the discovery/precursor step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly lists triggering user phrasings ('I have a folder of specs/files, what's in it?' or 'mock this directory'), which tells an agent when to invoke it. It also notes the one-level-only scan behavior, implying that deeper recursion requires further calls, though it does not explicitly name alternative sibling tools or when-not-to-use scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contextARead-only
Return the org, role and access (read or write) the current MCP credential is scoped to. Use this once at the start of a session to know which org you are acting in. A read connection cannot deploy; the user has to log in again with write access.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While the readOnlyHint annotation already marks this as safe, the description goes beyond it by explaining the practical consequence of read-scoped access: it cannot deploy and requires reauthentication. This gives the agent useful behavioral context that the annotation alone does not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: three sentences, each earning its place. The first sentence states the core purpose, the second gives a direct usage instruction, and the third explains the access limitation and required remediation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with no output schema, the description fully equips the agent: it names what is returned, when to call it, and the consequence of read access. No additional context is necessary for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema has full coverage, so there is no parameter ambiguity to resolve. The description appropriately focuses on the return values instead, listing org, role, and access, which is the relevant semantic information here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'Return', and names the exact resource: the org, role, and access scoped to the current MCP credential. This clearly distinguishes it from sibling tools like login, logout, and info by focusing on credential context rather than authentication or generic information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs the agent to use this once at the start of a session to know which org it is acting in. It also provides a key exclusion: a read connection cannot deploy, and the user must log in again with write access, helping the agent decide when to call login instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infoA
Summarise an OpenAPI spec or .mockz package without serving it. For a spec, returns {title, version, openapi_version, endpoint_count, paths} with operation IDs. For a package, returns its manifest. Pass input as a local file or a public https URL. Use this when the user wants to know what's in a spec or package before deciding whether to serve or deploy it.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | Spec file path, public https URL, or .mockz package. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it delivers: it discloses the non-serving read-only nature, the exact return shape for specs ({title, version, openapi_version, endpoint_count, paths} with operation IDs), the manifest return for packages, and accepted input forms. It does not cover error behavior or edge cases, which keeps it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five short sentences, each earning its place: purpose/scope, spec output shape, package output shape, input format, and usage trigger. The verb+scope is front-loaded and there is zero redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must disclose both safety and return values — and it does, enumerating outputs for both input kinds and framing the tool as read-only. The lone gap is a slight ambiguity: the sentence 'Pass input as a local file or a public https URL' follows the package clause without scoping, leaving it unclear whether packages also accept URLs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds modest clarification by phrasing accepted input as 'a local file or a public https URL' and qualifying the URL as public, but this mostly reinforces rather than extends the schema. No meaningful new semantics are added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Summarise an OpenAPI spec or .mockz package') and immediately distinguishes itself from the serving/deploying siblings with 'without serving it'. Among tools like serve_locally, deploy_mock_from_spec, and deploy_mock_from_url, an agent can tell this apart without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the trigger condition: 'Use this when the user wants to know what's in a spec or package before deciding whether to serve or deploy it.' The qualifier 'without serving it' also implies the exclusion of serving/deploying tools. It falls just short of the top bar because no specific sibling tool is named as the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
install_cliA
Install the mockzilla CLI for this user. Three methods — ASK the user which one they want before calling:
• download (recommended): fetch the prebuilt binary for this OS/arch from github.com/mockzilla/mockzilla releases (~38MB). Fast, no toolchain needed.
• go-install: run go install <module>@v<version> to compile from source. Needs Go on PATH.
• go-run: don't install at all — the bridge stores a go run <module>@v<version> invocation. First serve_locally compiles into Go's module cache; later runs are instant. Needs Go.
Files land in the bridge's own cache, never on system PATH; blow it away with rm -rf ~/.cache/mockzilla-mcp.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | download |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses installation methods, cache location, system PATH impact, and uninstall procedure. It is transparent about where files land and how to clean up, though could mention potential network usage or error scenarios.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (4 sentences) and front-loads the purpose and a key instruction. It could be improved with bullet points for the methods, but remains clear and avoids redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, no output schema), the description is complete: it explains each method, side effects (cache location, not on PATH), and how to uninstall. No critical gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds substantial meaning beyond the schema: it explains each enum value (download: prebuilt binary, go-install: compile from source, go-run: use go run without install), including size, speed, and dependencies. This is essential given 0% schema description coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Install the mockzilla CLI for this user' and details three specific methods, distinguishing the tool from siblings like check_cli and serve_locally.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description instructs the agent to ask the user which method to use and lists prerequisites for each method (e.g., Go needed for go-install/go-run). It does not explicitly exclude scenarios but provides sufficient context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lintA
Check an OpenAPI spec for schemas no value can satisfy, such as an array with a scalar enum or additionalProperties: false next to oneOf properties. Mockzilla can't generate valid responses for these, so requests to affected endpoints fail validation. Returns {clean, defect_count, defects: [{rule, path, detail}], truncated}; at most 50 defects are listed. Pass input as a local spec file or a public https URL. Use it before serving or deploying a spec, or when a mock returns unexpected validation errors. OpenAPI 3.x only.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | Spec file path or public https URL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely succeeds: it discloses the return object structure, the 50-defect cap/truncation behavior, and the downstream consequence of unsatisfiable schemas. It does not explicitly state that the operation is non-mutating, but 'Check' and the return contract strongly imply a read-only analysis.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences deliver the purpose, defect examples, return contract, parameter format, usage timing, and version constraint without filler. The purpose is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter lint tool with no output schema, the description is complete: it covers the input type, return fields, result truncation, version restriction, and the situations that warrant calling it. Nothing essential to invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes `input` fully as 'Spec file path or public https URL,' and the description only restates this. It adds no new syntax, resolution, or error-handling details beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action ('Check') on a concrete resource ('an OpenAPI spec') with a precise goal: find schemas no value can satisfy. Concrete examples like 'array with a scalar enum' and the 'OpenAPI 3.x only' constraint make the tool's role unambiguous among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to invoke it: before serving/deploying a spec or when a mock returns unexpected validation errors. It does not name alternatives or state exclusions, but the timing guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_catalog_productsARead-only
List the catalog specs (Stripe, Adyen, etc.) available to attach to a HOSTED sim on mockzilla.org. Returns recommended runtime settings the agent should compare against the org's tier before suggesting a deploy. Pass search to substring-match by slug or name. The returned ids/slugs are ONLY usable with deploy_mock_from_catalog — they are NOT URLs and NOT valid input for serve_locally. If the user wants to try a catalog product locally instead, skip this tool and call serve_locally with the public OpenAPI URL for the service (recall it from your training knowledge — Stripe, Twilio, etc. all publish OpenAPI specs on GitHub).
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| limit | No | ||
| search | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds important non-obvious behavior: it returns recommended runtime settings to compare against the org's tier, and the returned ids/slugs are ONLY usable with `deploy_mock_from_catalog` — not URLs and not valid for `serve_locally`. This materially helps the agent avoid misusing results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the main purpose, then adds return behavior, search guidance, and critical usage constraints, then closes with the local alternative. Every sentence adds necessary routing or behavioral information; there is no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the readOnly annotation, no output schema, and three optional parameters, the description covers everything an agent needs: purpose, result semantics, parameter hints, cross-tool constraints, and a local-workflow alternative. It is complete for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates by explaining `search` precisely: substring-match by slug or name. `page` and `limit` are left to their conventional pagination meanings and schema constraints, which are reasonably self-evident. The key domain-specific parameter, `search`, is fully clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List the catalog specs... available to attach to a HOSTED sim on mockzilla.org.' It also distinguishes this tool from siblings like serve_locally and deploy_mock_from_catalog by explaining what the returned IDs are and are not for. This leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage conditions: use `search` for substring matching, use this tool for hosted deploys, and skip it in favor of `serve_locally` when the user wants a local catalog product. It even names the alternative tool and the exact input format for that fallback, making the decision path explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_github_reposA
List the user's GitHub repositories so you can ASK which one to publish mocks to. Never pick one yourself. publishes_mocks: true means that repo already has the Mockzilla workflow, so publishing there updates its existing mock instead of creating another. Useful in clients with no shell, where you cannot run gh repo list.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| owner | No | Limit to a user or org. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and adds useful context: it explains the meaning of the `publishes_mocks: true` flag and the consequence for publishing (updates existing mock vs creating another). It does not detail return format or auth requirements, but for a simple read-only list operation the disclosed behavior is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each earning its place: purpose, action rule, flag meaning, and environment context. The key instruction is front-loaded, and there is no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-optional-parameter list tool with no output schema, the description covers the purpose, user interaction requirement, key flag semantics, and when it is most useful. Minor gaps like pagination behavior or auth prerequisites are not critical enough to lower further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not explain either parameter. `owner` has a schema description, but `limit` is only constrained by default/min/max with no stated meaning (e.g., number of repos to return). At 50% schema description coverage, the description should compensate for the undocumented parameter, and it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the user's GitHub repositories') and immediately clarifies the purpose: to ask the user which repo to publish mocks to. It distinguishes itself from publish-related siblings by emphasizing it is only a listing and decision-support step, not an action that publishes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context on when to use the tool ('clients with no shell, where you cannot run `gh repo list`') and gives an explicit behavioral rule ('Never pick one yourself'). It does not name an alternative sibling tool or state explicit when-not-to-use conditions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mock_endpointsA
List all endpoints currently mocked via mock_endpoint. Returns {endpoints: [{method, service, path, status, headers?, file}], server_url, ui_url}. An endpoint with body: null answers with no body. If a managed server is running, ui_url is the mockzilla UI (opens in a browser, shows endpoints grouped by service plus request inspection). Suggest the UI to the user when they want to explore beyond what the agent can show in chat.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It details the return structure, notes the condition for ui_url (only when a managed server is running), and explains the body:null behavior. This goes beyond a simple 'lists endpoints' statement, though it doesn't mention any potential errors or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no wasted words. The first sentence states the core purpose, the second provides essential return details, and the third adds a practical usage pointer. It is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with no parameters and no output schema, the description fully covers what an agent needs: the exact return format, conditional UI behavior, and a situational usage hint. Nothing is missing that would prevent correct invocation or interpretation of results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema is empty (coverage 100%). There is nothing to describe beyond the schema, so the baseline of 4 for parameterless tools applies. The description doesn't need to add parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the exact verb and resource: 'List all endpoints currently mocked via mock_endpoint.' It clearly distinguishes the tool from siblings like clear_mock_endpoints and mock_endpoint by focusing on listing existing mocks. The scope is precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear context for when to use the tool (to see mocked endpoints) and adds a specific usage suggestion: 'Suggest the UI to the user when they want to explore beyond what the agent can show in chat.' However, it doesn't explicitly compare against alternative tools or state when not to use it, though the purpose makes the primary use obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_replaysA
List the replay recordings a service currently holds, so you can tell the user what is pinned and what a matching request will return. Use it after setup_replay to confirm a recording was captured, or when a response looks stale and you suspect it is being replayed rather than produced fresh. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | One recording's key, for its full detail. | |
| service | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations supplied, the description carries the behavioral disclosure burden. It states 'Read-only' and explains what the call reveals (pinned recordings, what a matching request returns). It does not cover auth, rate limits, or failure modes, but this is reasonable for a read-only list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the primary purpose comes first, followed by concrete use cases and a read-only note. Every sentence adds value and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema and annotations, the description adequately explains what the tool does, when to use it, and what results conceptually represent. Minor gaps exist around output format or errors, but the agent has enough context to invoke and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% because 'service' has no description in the schema. The description adds context that the service owns the recordings and that a key provides full detail, but it does little to define the service parameter's format or required semantics beyond that. The key parameter is already documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('List the replay recordings a service currently holds') and clarifies the tool's role in exposing pinned recordings and replay behavior. It distinguishes itself from siblings like request_history and setup_replay by focusing on recording inventory rather than history or setup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger scenarios: use after setup_replay to confirm capture, or when suspicious that a response is replayed. This is clear contextual guidance, though it does not explicitly name alternatives or exclusions for when not to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_simsARead-only
List the sims (deployed mocks) accessible to the current org. Returns a page of sim entries with their refs, statuses, and live URLs. Pass search to substring-filter by name or sim_pk, or sim to look up exactly one sim_pk.
| Name | Required | Description | Default |
|---|---|---|---|
| sim | No | ||
| page | No | ||
| limit | No | ||
| search | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safety profile whether this is a safe read operation. The description adds meaningful behavioral context by describing pagination, output contents, and the filtering modes, without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first states purpose and output shape, the second gives the two filtering modes. No repetition, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list operation with readOnlyHint and no output schema, the description conveys scope, return contents, and filtering semantics. It is not fully complete because pagination parameters are not spelled out in prose, though they are visible in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clearly defines 'sim' and 'search', but leaves 'page' and 'limit' to be inferred from the phrase 'Returns a page.' That is helpful but not fully compensating for the absence of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('List the sims (deployed mocks)') and scopes them to the current org. It also clarifies the returned fields (refs, statuses, live URLs), making it semantically distinct from sibling tools like list_mock_endpoints even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear operational context: 'accessible to the current org' and 'Returns a page of sim entries.' It also gives explicit parameter usage rules for 'search' and 'sim', but it does not contrast this tool with alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
loginA
Log in to Mockzilla cloud so the hosted tools (deploying mocks, listing simulations, the catalog) become available. Opens the Mockzilla login in the user's browser, where they pick an organization and read-only or read-and-write access. Returns right away with the login url. Call it when a hosted tool says it needs a login, then call that tool again once the user approves. Local tools never need it. Side effects: starts a short-lived listener on 127.0.0.1 for the login callback, and saves the login under ~/.config/mockzilla-mcp/.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does so well: it discloses browser opening, the short-lived 127.0.0.1 listener, the config file location, and that it returns immediately with a URL. This gives the agent a realistic model of the side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact yet information-dense, starting with the main purpose and ending with side effects. Every sentence contributes: purpose, flow, return behavior, when to use, and side effects.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter login tool with no output schema, this is complete: the agent knows what triggers login, what happens in the browser, what is returned, where credentials are stored, and how to retry the originating tool. No critical operational detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema is empty, so the baseline is 4. The description adds useful context by explaining the browser interaction and access-choice flow, which is more than the schema alone could convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Log in to Mockzilla cloud') and the concrete effect: hosted tools become available. It also distinguishes itself from local tools and from logout by naming the hosted-tool context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to call it ('when a hosted tool says it needs a login') and what to do next ('call that tool again once the user approves'). It also says local tools never need it, giving a clear exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
logoutA
Log out of Mockzilla cloud on this machine. Revokes the connection and deletes the saved login; hosted tools disappear until the next login. Call it when the user asks to log out or to switch organization or access level.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool revokes the connection, deletes saved login, and causes hosted tools to disappear until next login. This is meaningful behavioral context beyond a simple 'log out' statement. It doesn't mention side effects like whether other machines are affected, but for a logout tool this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the action, the behavioral effects, and the usage trigger. Front-loaded with the verb and resource. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter logout tool with no output schema, the description covers the action, effects, and when to use it. It could mention whether the user will be prompted for confirmation or whether it affects other sessions, but those are minor gaps. The description is complete enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is trivially complete. The description adds context about what the operation does, which is more than needed. Baseline 4 for zero params is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Log out' and the resource 'Mockzilla cloud on this machine', and distinguishes it from the sibling 'login' by describing the effect (revokes connection, deletes saved login). It also mentions hosted tools disappearing until next login, which adds clarity about the scope of the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to call it: 'when the user asks to log out or to switch organization or access level.' This provides clear usage context and implicitly distinguishes it from login and other tools. It doesn't name alternatives explicitly, but the condition is specific enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mock_endpointA
Quickly mock a single HTTP endpoint without writing an OpenAPI spec. Pass method (default GET), path (the EXACT HTTP path the user described, including all segments), and the response body (object → JSON, string → text). The bridge writes the response into a managed static dir at ~/.cache/mockzilla-mcp/mocks/ and (re)starts a single shared mockzilla server pointing at it.
Pass path AS IS. Do NOT prepend or duplicate any segment. The bridge derives the service name from the first segment for internal grouping, but it does not change the URL the user hits. Examples:
• User says GET /pets/{id} → call mock_endpoint with path=/pets/{id} → URL is http://HOST:PORT/pets/{id}
• User says POST /orders → path=/orders → URL is http://HOST:PORT/orders
• User says GET /v1/users/me → path=/v1/users/me → URL is http://HOST:PORT/v1/users/me
Pass status and/or headers to mock a failure or a redirect with a real body: status: 404 with an error payload, 201 with a Location, 429 with Retry-After. Omit response to send no body at all, which 204 and 304 require. These need mockzilla 2.8.20 or newer; the tool says so if the installed CLI is older. To fail a share of requests instead of every one, use serve_locally with errors.
Path placeholders like {id} are stored as literal directory names — by default ALL placeholder values share the same response. To return different responses for specific values, call mock_endpoint again with a literal value (e.g. /pets/123).
Calling this multiple times accumulates endpoints in the same server — adding POST /pets after GET /pets/{id} keeps both. Mutually exclusive with serve_locally: stop any ad-hoc server first. See mockzilla_docs_search('static directory') for the underlying convention.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path beginning with /. First segment is the service name. | |
| method | No | GET | |
| status | No | Response status. Omit `response` alongside it for a body-less response, which is what 204 and 304 need. | |
| headers | No | Extra response headers, e.g. {"Location": "/orders/42"}. A Content-Type here wins over `content_type`. | |
| response | No | Response body. Object → JSON. String → text. Default {}. | |
| content_type | No | Override content type. Inferred from response type if omitted (object → application/json, string → text/plain). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does so thoroughly. It discloses the write location ~/.cache/mockzilla-mcp/mocks/, server (re)start behavior, endpoint accumulation across calls, literal storage of placeholders, version requirements, and side effects of repeated calls. This is exemplary behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but tightly organized: primary use case first, then path rules with examples, then status/header behaviors, then accumulation and conflict rules. Each sentence earns its place by adding operational detail rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers selection, invocation semantics, side effects, version caveats, and interactions with related tools. An agent has everything needed to call it correctly and predict the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even with 83% schema coverage, the description adds substantial value: path must be passed exactly as-is with explicit examples, response body conversion depends on type, status/headers enable failure and redirect responses, and omitting response is required for 204/304. These details go well beyond the schema property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states exactly what the tool does: 'Quickly mock a single HTTP endpoint without writing an OpenAPI spec.' This clearly distinguishes it from sibling tools like deploy_mock_from_spec and serve_locally. The examples reinforce the precise scope with concrete call patterns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: quick single endpoint mocking without a spec. It also names an alternative for a different need—'use serve_locally with errors' for partial failures—and states the mutual exclusivity with serve_locally. This is strong routing that an agent can act on immediately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mockzilla_docs_readA
Return the full markdown of Mockzilla doc topics. Pass topics with one or more ids from mockzilla_docs_topics, or category to read a whole category in one call. A large request returns what fits and lists the rest in remaining: read those in a second call. A link to another topic reads as (topic id). Each topic carries url, its public page, to give the user; don't fetch it, the markdown is the same page.
| Name | Required | Description | Default |
|---|---|---|---|
| topics | No | Topic ids from `mockzilla_docs_topics`, e.g. ['getting-started/terminology', 'engine/contexts']. | |
| category | No | A category slug from `mockzilla_docs_topics`, e.g. 'simulations'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals truncation behavior with `remaining`, explains how inter-topic links render as `(topic id)`, and clarifies that each topic includes a public `url` that should be given to the user but not fetched. This is detailed and useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, all substantive and front-loaded with the core purpose. Every clause adds operational detail—truncation, link handling, and URL usage—without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no output schema and no annotations, the description is exceptionally complete. It covers input selection, response truncation and pagination, link representation, and user-facing URL behavior. An agent has what it needs to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters, so the baseline is 3. The description adds meaningful semantics: `category` reads a whole category in one call, `topics` accepts one or more ids, and large responses may require follow-up reads via `remaining`. This goes beyond the schema's basic type hints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return'), the object ('full markdown'), and the scope ('Mockzilla doc topics'). It is distinct from siblings like mockzilla_docs_topics and mockzilla_docs_search, making it easy for an agent to know this tool reads full doc content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete direction on when to use `topics` vs `category`, and explicitly tells the agent not to fetch the `url` because the markdown is the same page. It references mockzilla_docs_topics for id sourcing, though it does not explicitly contrast this tool with mockzilla_docs_search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mockzilla_docs_searchA
Search the Mockzilla docs by keyword. Returns the best matching sections {topic, title, heading, snippet} so you know which topics to read. Use it when no topic title from mockzilla_docs_topics clearly fits. Answer from the docs rather than from memory: they describe the product as it is now.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | Free-text query, e.g. 'static directory layout'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It clearly states this is a search operation that returns best-matching sections and gives the exact fields returned. It also notes the docs describe the current product, which is useful context. It doesn't mention edge cases like empty results, but for a read-only docs search this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, each earning its place: what the tool does, what it returns, and when to use it. The key information is front-loaded in the first sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with two parameters and no output schema, the description covers purpose, return shape, and usage context. It could be slightly stronger by explicitly pointing to mockzilla_docs_read for fetching full sections, but the phrase 'so you know which topics to read' implies the next step.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the query parameter with an example. The description reinforces that search is by keyword but adds no new meaning. The limit parameter is not mentioned in the description, though its name, default, and bounds in the schema make its role reasonably inferable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Search the Mockzilla docs by keyword') and defines the return shape ({topic, title, heading, snippet}). It also differentiates itself from mockzilla_docs_topics by explicitly naming when that sibling should be preferred, so an agent can distinguish the tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use rule: use this when no topic title from mockzilla_docs_topics clearly fits. It also gives a behavioral directive to answer from docs rather than memory, which clarifies the tool's role in the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mockzilla_docs_topicsA
List the Mockzilla docs that ship with this bridge: the product docs from mockzilla.org (what a simulation is, deploying, resilient backends, billing, settings, the CLI and this MCP server) and the open-source engine docs under engine/ (configuration, contexts, middleware, replay). Returns each category with its topics' id, title and one-line summary. The docs are files inside the bridge, so this needs no network and no login. Call it before answering a question about Mockzilla, then read the topics that fit.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure. It states that the docs are files inside the bridge, requiring no network or login, and describes the output format. This covers the key behavioral traits—read-only, offline, and deterministic. Some minor aspects like potential size limits or pagination are not mentioned, but for a simple listing tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose, followed by output details and a usage tip. Every sentence earns its place: purpose, content breakdown, key behaviors, and when to call it. No redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description is fully complete. It specifies what the tool returns (categories with id, title, summary), the content scope, the operational context (no network/login), and even gives usage guidance. An agent has everything needed to invoke and interpret the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100%, so the description adds no parameter-specific meaning. According to the rubric, a 0-parameter tool gets a baseline of 4. The description's mention of output categories doesn't compensate for missing params (none exist), so it remains at baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific verb 'List' and resource 'Mockzilla docs', and distinguishes itself from siblings like mockzilla_docs_read and mockzilla_docs_search by describing its unique output (categories with topics' id, title, summary). It also specifies the scope (product docs and engine docs), making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs 'Call it before answering a question about Mockzilla, then read the topics that fit,' which provides a clear when-to-use directive. It implies a follow-up step (reading topics) without explicitly naming the read tool, but the context is unambiguous. It doesn't state exclusions (when not to use), but the positive guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
packA
Pack a directory of mockzilla services into a .mockz archive for easier distribution or sharing. The archive carries a manifest (name, description, mounts, modes, git source) so the runtime can register every service without re-walking the tree. Hand the resulting .mockz to anyone — they can serve it with serve_locally, even from a URL. Use this when the user wants to share a working mock setup, snapshot one for a teammate, or publish it.
Defaults: output is <basename>.mockz next to dir. Git metadata (remote, ref, commit) is auto-embedded when dir is inside a git tree — pass skip_git: true to suppress.
| Name | Required | Description | Default |
|---|---|---|---|
| dir | Yes | Directory containing services / specs to pack. | |
| name | No | Display name embedded in the archive manifest. | |
| output | No | Output .mockz path. Defaults to <basename>.mockz next to dir. | |
| skip_git | No | Don't auto-embed git remote/ref/commit in the manifest. | |
| description | No | Free-text description embedded in the manifest. | |
| min_version | No | Minimum mockzilla version required to load this archive (e.g. '2.5.3'). Useful when the archive relies on newer features. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It does well: it reveals the manifest contents, the default output location, and the auto-embedding of git metadata with a suppression flag. It doesn't discuss overwrite behavior or what is returned, leaving minor gaps, but the core side effects and defaults are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: first the action/result, then when to use it, then defaults and edge behavior. Every sentence adds information, with no filler or repetition of structured schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema and no annotations, the description covers the key call decisions: required dir, output default, git embedding, and intended usage. It doesn't explicitly state the return value or overwrite policy, but the archive-creation behavior and consumer workflow are sufficiently described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning by explaining the conditional git-metadata behavior (embedded only when dir is inside a git tree) and how skip_git suppresses it, which goes slightly beyond the bare schema. Other parameter defaults are also reiterated and reinforced in context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names the action (pack), the input (directory of mockzilla services), and the concrete output (.mockz archive). It unambiguously differentiates this from sibling deploy/search tools by establishing that pack produces a distributable bundle rather than starting or serving anything.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to reach for this tool: sharing a mock setup, snapshotting for a teammate, or publishing. It also explains how the result can be consumed via serve_locally. It does not list alternatives or when-not-to-use conditions, but the when-to-use context plus the consumer note is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publish_to_githubA
Publish mocks to one of the USER'S OWN GitHub repositories and let the Mockzilla action deploy them at a shareable host of their own, https://.api.mockz.io. The label is the repository name, with -2, -3 if it is taken. No Mockzilla account is needed: the first push registers the repository. This is the free path for a user who is not logged in, and a valid choice for one who is when they want the mocks living in a repo and reviewed like code. If they are logged in and just want a quick hosted mock, prefer deploy_mock_from_* instead, which also gives history and replays in the app.
ASK THE USER FIRST, do not guess:
WHICH REPOSITORY. It is theirs, not one you invent. It can be an existing repo, including an app repo they already have, since this only adds a services folder and a workflow.
list_github_reposshows the candidates.PRIVATE OR PUBLIC, if it has to be created.
visibilityis required and has no default. The deployed mock URL is public either way, so say so: whatever is in these responses is readable by anyone with the link.
SIDE EFFECTS: uses the user's own gh login, may create a repository, commits and pushes, and triggers a public deploy. An existing services folder is merged into, not replaced, unless replace is true; an existing workflow is never overwritten.
Mocks here are static: spec-generated or fixed responses. If the user wants real logic or state, this is the wrong tool; that needs the codegen action and a Go server, and this refuses to publish into such a repository. Keep specs small too, since a free simulation has 128MB and a big spec costs far more in memory than on disk: simplify first if needed. Then call wait_for_github_deploy. The host is picked when the deploy runs, so this tool returns no URL and that one does.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | "owner/name" of the user's repo. Created if absent. | |
| source | No | Folder to publish. Defaults to the mocks built by mock_endpoint. | |
| message | No | Commit message. | |
| replace | No | Delete the repo's existing service folder first. Default merges. | |
| visibility | Yes | Required when creating. Ask the user; there is no default. | |
| description | No | Description for a new repo. | |
| services_dir | No | Where service folders go in the repo. Defaults to "services"; use another path when adding mocks to an existing project. | services |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure and does so strongly. It states side effects up front: uses the user's gh login, may create a repository, commits and pushes, triggers a public deploy, merges an existing services folder unless replace is true, and never overwrites an existing workflow. It also reveals that the public URL means responses are readable by anyone with the link.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and organized into clear paragraphs for usage, required user decisions, side effects, and constraints. It is long, but the complexity of the tool justifies most of it; a few phrases are slightly redundant with the schema, such as visibility being required with no default.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, externally visible action with seven parameters and no output schema, the description is remarkably complete. It covers prerequisites, user confirmation steps, side effects, what the tool refuses to do, spec-size limits, and the follow-up tool to call. It even explains the absence of a returned URL, which is essential when there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already handles basic parameter definitions. The description adds worthwhile operational meaning beyond the schema: repo must be user-owned and not invented, candidates can come from list_github_repos, visibility has no default and the resulting URL is public either way, and replace/services_dir merge behavior is clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: publishing mocks to the user's own GitHub repository and letting the Mockzilla action deploy them. It also differentiates from sibling tools by labeling this the free path for unauthenticated users, a code-review-friendly option, and by explicitly noting that this tool returns no URL while wait_for_github_deploy does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-to-use and when-not-to-use guidance: prefer deploy_mock_from_* for quick logged-in hosted mocks, use this for repo-based/free hosting, and avoid it for real logic or state, which needs the codegen path. It also chains the workflow by instructing the agent to ask the user first, consult list_github_repos, possibly run simplify, and then call wait_for_github_deploy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_historyA
List the requests the running local server has answered, newest first: method, URL, status, content type, how long it took, and where the response came from (generated, upstream, replay or cache). Use it to show the user what their code actually sent, to confirm a call arrived, or to find the request to look at next. Pass id together with service for one request's full headers and body. Reads the server's own history API, so it needs no account and works entirely on localhost. To ask what is WRONG with the traffic rather than list it, call diagnose_requests.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | One entry's id, from a previous listing. Needs `service` too. Returns full request and response headers and bodies. | |
| limit | No | ||
| method | No | Only this HTTP method, e.g. "POST". | |
| status | No | Only this exact status code. | |
| service | No | Limit to one service. Omit to read every service the server has. | |
| failed_only | No | Only responses with status 400 or above. | |
| path_contains | No | Only requests whose URL contains this substring. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden, and it does well: it characterizes the tool as read-only ('Reads the server's own history API'), clarifies auth needs ('needs no account'), and scopes it to localhost. It does not go into edge-case behavior such as id-without-service or whether listing mutates state, but the explicit read framing largely covers the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence serves a purpose: what the tool returns, when to use it, how to get detail, its auth/localhost behavior, and when to pick a sibling. The main result is front-loaded and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 optional parameters and no output schema, the description does enough: it names the returned fields, explains the id/service detail mode, gives the local/no-account context, and routes to diagnose_requests. It does not specify how filters interact with id or describe pagination beyond 'newest first', but the sibling routing and core usage guidance cover the main agent needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already explains most parameters. The description adds context for the id+service pair and explains what the returned history contains, but it does not add meaning beyond the schema's parameter descriptions. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('List the requests the running local server has answered'), states the result ordering, and enumerates the returned fields. It also distinguishes itself from sibling diagnose_requests by explicitly saying that tool is for asking what is WRONG with traffic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete use cases: show what the user's code sent, confirm a call arrived, and find the request to look at next. It also provides an explicit exclusion by directing agents to diagnose_requests when they want problem analysis rather than a listing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serve_locallyA
Start ONE mockzilla portable mock server on this machine that serves any number of APIs together — no mockzilla account needed. Pass input as a single spec path / directory / public https URL, OR an array of them to combine multiple APIs into the same server (each becomes a service mounted at //...). Returns {url, port, pid, services} plus example_endpoints, callable URLs for a single spec. Use one of those rather than guessing a path: each service answers under its mount prefix, not at the spec's bare path. Pair with stop_locally(pid) to clean up. Prefer this over deploy_mock_from_* whenever the user says 'try locally', 'experiment', or 'play with' — those tools create persistent hosted bundles, this one is ephemeral. The bridge only runs ONE local server at a time on purpose: if the user wants more APIs, stop the current server and restart with all of them in input.
To test how a client handles a slow or failing API, pass latency and/or errors. These and mount/context work only when input is a single spec or single-service folder.
If the user names a well-known API (stripe, twilio, github, openai, slack, etc.) WITHOUT providing a URL, recall the public OpenAPI spec URL from your training knowledge and pass that. Do NOT pass a catalog ID or slug from list_catalog_products — that catalog is for the HOSTED deploy_mock_from_catalog flow, its ids are not URLs. Examples of public OpenAPI URLs:
• Stripe: https://raw.githubusercontent.com/stripe/openapi/master/openapi/spec3.json
• Twilio: https://raw.githubusercontent.com/twilio/twilio-oai/main/spec/json/twilio_api_v2010.json
• GitHub: https://raw.githubusercontent.com/github/rest-api-description/main/descriptions/api.github.com/api.github.com.json
• Petstore: https://petstore3.swagger.io/api/v3/openapi.json
| Name | Required | Description | Default |
|---|---|---|---|
| port | No | Port to bind on. Omit or pass 0 to let the OS pick a free port. | |
| input | Yes | Spec file path(s), directory, or public OpenAPI URL(s). Pass an array to combine multiple APIs into one server. | |
| mount | No | URL path to mount the service at, e.g. "pets/v2". | |
| errors | No | Error injection by cumulative percentile, keys p1 to p100. {"p5": 500, "p10": 503} returns 500 for 5% of requests and 503 for the next 5%. | |
| context | No | Path to a flat context YAML of replacement values for generated data. | |
| latency | No | Delay added to every response, as a Go duration: "100ms", "1.5s". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so thoroughly. It discloses ephemerality, single-server limitation, mount-prefix behavior, the fact that latency/errors/mount/context only work for single specs, and warns against passing catalog IDs. It even explains the returned fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed with essential usage constraints, examples, and warnings. It front-loads the core purpose and then layers behavioral rules and alternatives. Slightly verbose with the URL examples, but each part earns its place for this complex tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, 6 parameters, no output schema, and many sibling alternatives, the description covers everything an agent needs: return shape, constraints, input types, known-API handling, single-server behavior, and cleanup. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by clarifying that input can be a single path/URL or an array, that latency/errors/mount/context apply only to single-spec inputs, and that well-known API names should be resolved to public OpenAPI URLs before passing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool starts one local mockzilla server serving any number of APIs together, with explicit contrast to the deploy_mock_from_* siblings. It names the resource, action, and scope precisely, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to prefer this over deploy_mock_from_* for 'try locally', 'experiment', or 'play with', and explains those alternatives create persistent hosted bundles. It also instructs pairing with stop_locally(pid) and describes the one-server-at-a-time rule with restart guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_replayA
Configure replay for a service: record a real response once, then serve it back for every matching request (VCR). Writes a replay: block into the service's config.yml.
ASK THE USER TWO THINGS BEFORE CALLING:
Record from a real backend, or pin the mock's own output? With
upstream_urlthe recording is the real backend's response. Without it, replay pins the generated response so repeat calls stop returning fresh random data, which is often what 'make it stable' means.One recording for the whole endpoint, or one per input? With no
matchfields the key is method and path only, so EVERY call to POST /foo replays the first response no matter what it sends. Look at what the endpoint actually takes, then ask which fields distinguish one case from another and pass those asmatch.
Match fields come from three sources: path (path variables, ignored unless listed), body (dotted paths like data.items[0].name, or "[0].name" for a top-level array, or flat keys for form bodies), and query. Returns recording_scope spelling out what each endpoint is keyed by. Only writes inside the bridge's own mocks dir; for a folder served from the user's project it returns the YAML and the path for them to apply. Config is read at startup, so restart after.
| Name | Required | Description | Default |
|---|---|---|---|
| dir | No | Directory holding the service, when it is not one of the bridge's own mocks. | |
| service | Yes | Service to record, as named in serve_locally's output. | |
| duration | No | How long recordings live, e.g. "24h". | |
| endpoints | Yes | Endpoints to record. | |
| auto_replay | No | Record and replay without the X-Mockzilla-Replay header. | |
| upstream_url | No | Real backend to record from, e.g. "https://api.example.com". Omit to record the mock's own generated responses. | |
| upstream_only | No | Refuse to record anything that did not come from the upstream. Needs `upstream_url`, or every request answers 502. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers. It discloses the write side effect, the write scope ('only writes inside the bridge's own mocks dir'), the fallback behavior for user-served folders, and the requirement to restart because config is read at startup. This is exactly the kind of behavioral context agents need.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section earns its place: core concept, mandatory user questions, matching syntax, write location, and restart requirement. It is front-loaded with the essential behavior and structured so an agent can extract the call-critical information quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description tells the agent what the tool returns ('recording_scope spelling out what each endpoint is keyed by'), what side effects occur, and what postconditions apply. For a multi-option setup tool, this is complete enough to call correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining operational semantics: match fields come from path, body, and query; body supports dotted paths; path variables are ignored unless listed; and omitting match keys on method+path alone. This meaningfully exceeds what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: 'Configure replay for a service' and immediately defines the behavior with a VCR analogy: 'record a real response once, then serve it back for every matching request.' It also states the concrete side effect, writing a `replay:` block into config.yml, which distinguishes it clearly from sibling tools like list_replays.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit pre-call guidance by telling the agent to ask the user two questions before calling, and explains the practical consequences of each choice. It doesn't name alternative tools or exclusions, but the decision context it provides is strong enough for an agent to know when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
simplifyA
Simplify an OpenAPI spec: drop or reduce union types (anyOf/oneOf), strip x-* extensions, and optionally limit the number of optional properties per schema. Writes the simplified spec to disk and returns its path. Use this when a spec is too large or too complex to mock cleanly (deeply nested unions, hundreds of optional fields) — the output is a faithful subset the agent can hand to serve_locally.
Optional-property handling:
• omit optional to keep every optional property
• optional: N keeps exactly N per schema (0 drops them all)
• optional_min/optional_max (must come together) picks a random count in that range per schema
Pass config for an oapi-codegen-dd codegen.yml when the user wants filter + overlay + prune applied before simplification.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | Spec file path or public https URL. | |
| config | No | Path to an oapi-codegen-dd codegen.yml. Applies filter + overlay + prune before simplification. | |
| output | No | Where to write the simplified spec. Omit to write into ~/.cache/mockzilla-mcp/simplified/ and let the bridge pick a name. | |
| optional | No | Keep exactly N optional properties per schema. Mutually exclusive with optional_min/optional_max. | |
| optional_max | No | Range mode upper bound (use with optional_min). | |
| optional_min | No | Range mode lower bound (use with optional_max). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden—and it delivers. It discloses that the tool writes to disk and returns the path, that `config` causes a pre-processing pipeline (filter + overlay + prune), and that the result is 'a faithful subset.' It does not mention side effects like overwriting existing files or how the random count is seeded, so it's not a perfect 5, but the main behavioral surface is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first two sentences, followed by a compact bulleted list for the optional-property parameter modes. Every sentence adds information—there is no filler or restatement of the input schema. The structure makes the trickiest parameter semantics scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six parameters and no output schema, this description covers the input source, the transformation behavior, the output destination fallback, the full parameter-mode matrix, and the `config` pipeline. The only minor omissions (overwrite behavior, exact return-value shape) are understandable given the description text and the 100% schema coverage; nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3; the description adds genuine value on top by explaining the semantics of `optional` modes (keep all vs. exactly N vs. random range), the mutual exclusivity of `optional` with `optional_min`/`optional_max`, and the `config` preprocessing pipeline. It loses one point because `output`'s fallback behavior is stated but the random-range semantics could have been slightly clearer about the inclusive/exclusive bounds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (simplify) and resource (OpenAPI spec), enumerates exactly what transformation is applied (drop/reduce anyOf/oneOf, strip x-* extensions, limit optional properties), and distinguishes its output use case from siblings by naming `serve_locally`. This is clearly differentiated from sibling tools like `lint`, `pack`, and `serve_locally`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool ('when a spec is too large or too complex to mock cleanly') and names the downstream sibling (`serve_locally`) that should receive the output. It also includes detailed mode-selection guidance for the optional-property parameters (omit, exact N, or min/max range), which routes the agent to the correct parameter combination and even notes the mutual-exclusion requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_locallyA
Stop the mockzilla server started by serve_locally. Takes no arguments — there's only ever one local server running. Returns {stopped: bool, pid?, reason?}.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description fully discloses behavior: stops server, takes no arguments, returns {stopped: bool, pid?, reason?}. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, each providing essential information. Action is front-loaded. No extraneous text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description explicitly states return format and server uniqueness, making it complete for agent usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters, so baseline is 4. Description adds context that it takes no arguments and only one server exists, which is helpful beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool stops the mockzilla server started by serve_locally, specifying the action and resource. Distinguishes from sibling tools by explicitly mentioning the server it stops.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage context: use after serve_locally when a local server is running. Notes there is only one server. Does not explicitly state when not to use, but the simplicity of the tool makes it adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unpublish_from_githubA
Take down the mocks a repository serves, freeing its simulation slot. Runs the workflow with delete: true, which needs the workflow_dispatch trigger this bridge writes; a repo whose own workflow lacks it cannot be torn down this way, and the result says so. Pass delete_repo: true to also delete the repository itself, which is permanent and needs the delete_repo scope, so confirm with the user before using it.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | ||
| delete_repo | No | Also delete the repository. Permanent. Ask first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does so thoroughly. It discloses the mechanism (runs workflow with delete: true), the prerequisite (workflow_dispatch trigger), the failure behavior ('the result says so'), the permanence of delete_repo, the required scope, and the need for user confirmation. This is comprehensive behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loads the core purpose, then explains mechanics and cautions. Every sentence adds essential information with no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, the description covers purpose, operation, prerequisites, side effects, and required permissions. It does not specify what the repo parameter should look like or describe the success/error return format, but the phrase 'the result says so' hints at error messaging. Overall it is near-complete for an agent to act correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (delete_repo has a schema description, repo does not). The description adds value by explaining delete_repo's scope requirement and permanence, but it does not clarify the expected format for the repo parameter (e.g., 'owner/repo'), leaving that gap. It partially compensates but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Take down the mocks a repository serves, freeing its simulation slot.' It uses a specific verb ('take down') and identifies the resource (repository), and this is the inverse of the sibling publish_to_github, making it distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool (to unpublish mocks) and explicitly states when it cannot work (if the repo's workflow lacks workflow_dispatch). It also instructs to confirm with the user before using delete_repo: true. However, it does not name alternative sibling tools or explicitly compare with publish_to_github, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_deployARead-only
Block until a sim's deploy reaches a terminal status (active or failed) or until timeout_s seconds elapse (default 25, max 30). Returns {sim_id, status, urls: {live, dashboard}} — urls.live is set once the deploy is ACTIVE. On timeout the response carries the current non-terminal status (typically deploying); the agent can call wait_for_deploy again with the same sim_id.
| Name | Required | Description | Default |
|---|---|---|---|
| sim_id | Yes | ||
| timeout_s | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses that the call blocks, how long it may block, what the return payload contains, which statuses are terminal, and what happens on timeout. It also tells the agent it can retry with the same sim_id, which is important behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: core blocking behavior first, then return shape, then timeout behavior and retry guidance. Every sentence earns its place and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and minimal annotations, the description provides the return shape, terminal statuses, timeout behavior, and a retry strategy. An agent has what it needs to invoke the tool correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It explains `timeout_s` with default and max values and its effect on blocking behavior, and it clarifies `sim_id` as the identifier of the sim whose deploy is being waited on, including reuse on retry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Block until') and names the exact resource and condition: a sim's deploy reaching `active` or `failed`. This unambiguously distinguishes it from deployment-creation and serving siblings like `deploy_mock_from_spec` and `serve_locally`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the operational context clear: call it to wait for a deploy to finish, and call it again on timeout. It does not explicitly state when not to use it or name alternative polling strategies, but the blocking behavior and retry guidance provide solid usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_github_deployA
Wait for the Mockzilla workflow run to finish and return the live mock URL, as the action printed it in the run's log. Mockzilla picks the host when the deploy runs, so this is the one place to get it. Call after publish_to_github, or after the user pushes. A first deploy takes a minute or two. If it returns with no conclusion it ran out of time: call again.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | ||
| timeout_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses that the tool reads the workflow log, waits for completion, and returns the URL; it also explains the timeout symptom ('returns with no conclusion') and tells the agent to call again. It does not cover auth or failure modes, but the core polling/log-reading behavior is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five short sentences, each earning its place. The primary action is front-loaded, and the description avoids filler while efficiently conveying sequencing, timing, and retry behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main orchestration context: when to call, how long to expect, what it returns, and what to do on timeout. It is slightly incomplete on the required repo argument and precise failure behavior, but given the tool's simplicity and no output schema, it provides enough for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions repo or timeout_seconds. The retry/timeout prose only indirectly relates to timeout_seconds, and repo is left entirely to inference from the tool name and GitHub context, so the description does not adequately compensate for the lack of schema-level parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation: wait for the Mockzilla workflow run to finish and return the live mock URL from the run log. The wording clearly identifies the GitHub/Mockzilla deployment context, making it easy to tell apart from sibling deploy and wait tools, though it never explicitly names wait_for_deploy as the competing alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing guidance: 'Call after publish_to_github, or after the user pushes' and explains the expected timing with 'A first deploy takes a minute or two.' It also provides a retry instruction for timeout cases, but it does not explicitly state when not to use this tool or name alternatives such as wait_for_deploy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.2.25- Added
check_github_deployable - Added
diagnose_requests - Added
list_github_repos - Added
list_replays - Added
publish_to_github - Added
request_history - Added
setup_replay - Added
unpublish_from_github - Added
wait_for_github_deploy
1 tool update
v0.2.22- Changed
mock_endpoint3 fields changed- added
Input schema / properties / headersAdded value: +{ + "additionalProperties": { + "type": "string" + }, + "description": "Extra response headers, e.g. {\"Location\": \"/orders/42\"}. A Content-Type here wins over `content_type`.", + "type": "object" +} - added
Input schema / properties / status / descriptionAdded value: +"Response status. Omit `response` alongside it for a body-less response, which is what 204 and 304 need." - changed
Input schema / properties / status / minimumPrevious value: -100New value: +200
16 tool updates
v0.2.20- Added
deploy_mock_from_catalog - Added
deploy_mock_from_spec - Added
deploy_mock_from_url - Added
get_context - Added
info - Added
lint - Added
list_catalog_products - Added
list_sims - Added
login - Added
logout - Changed
mockzilla_docs_read4 fields changed- added
Input schema / properties / categoryAdded value: +{ + "description": "A category slug from `mockzilla_docs_topics`, e.g. 'simulations'.", + "type": "string" +} - removed
Input schema / properties / topicRemoved value: -{ - "description": "Topic name from `mockzilla_docs_topics` (e.g. 'middleware', 'usage/portable').", - "type": "string" -} - added
Input schema / properties / topicsAdded value: +{ + "description": "Topic ids from `mockzilla_docs_topics`, e.g. ['getting-started/terminology', 'engine/contexts'].", + "items": { + "type": "string" + }, + "minItems": 1, + "type": "array" +} - removed
Input schema / requiredRemoved value: -[ - "topic" -]
- Added
pack - Removed
peek_openapi - Changed
serve_locally4 fields changed- added
Input schema / properties / contextAdded value: +{ + "description": "Path to a flat context YAML of replacement values for generated data.", + "type": "string" +} - added
Input schema / properties / errorsAdded value: +{ + "additionalProperties": { + "maximum": 599, + "minimum": 100, + "type": "integer" + }, + "description": "Error injection by cumulative percentile, keys p1 to p100. {\"p5\": 500, \"p10\": 503} returns 500 for 5% of requests and 503 for the next 5%.", + "type": "object" +} - added
Input schema / properties / latencyAdded value: +{ + "description": "Delay added to every response, as a Go duration: \"100ms\", \"1.5s\".", + "type": "string" +} - added
Input schema / properties / mountAdded value: +{ + "description": "URL path to mount the service at, e.g. \"pets/v2\".", + "type": "string" +}
- Added
simplify - Added
wait_for_deploy
14 tool updates
v0.1.0- First observed
bridge_status - First observed
call_endpoint - First observed
check_cli - First observed
clear_mock_endpoints - First observed
discover_specs - First observed
install_cli - First observed
list_mock_endpoints - First observed
mock_endpoint - First observed
mockzilla_docs_read - First observed
mockzilla_docs_search - First observed
mockzilla_docs_topics - First observed
peek_openapi - First observed
serve_locally - First observed
stop_locally
TDQS
Scored across 35 tools
Every tool has a clearly distinct purpose, and the descriptions actively cross-reference each other to prevent misselection: deploy_mock_from_* explicitly warns it is NOT valid input for serve_locally, request_history vs diagnose_requests are differentiated as list-vs-diagnose, and wait_for_deploy vs wait_for_github_deploy are separated by deployment target. Even the three near-identical deploy_mock_from_catalog/spec/url tools are cleanly distinguished by input source. With 35 tools this is an exceptional level of disambiguation.
The overwhelming majority follow a consistent snake_case verb_noun pattern (list_replays, serve_locally, deploy_mock_from_catalog, wait_for_github_deploy, unpublish_from_github), and related families share structural prefixes (mockzilla_docs_*, deploy_mock_from_*, wait_for_*). Minor deviations include bare verbs (info, lint, simplify, pack) and a noun-first name (bridge_status), but these are isolated and the overall pattern remains highly predictable.
35 tools is on the heavy side and exceeds the calibration's 25-tool threshold for 'too many', but the domain is genuinely broad: local serving, hosted deployment, GitHub publishing, replay, spec processing, CLI management, and docs each demand their own surface. Some consolidation is possible (deploy_mock_from_spec and deploy_mock_from_url could be one tool; wait_for_deploy vs wait_for_github_deploy) but each tool does earn a place in the platform's scope.
The surface covers the full lifecycle remarkably well: create (serve_locally, mock_endpoint, deploy_*), read (list_sims, list_mock_endpoints, request_history, info), modify (setup_replay, simplify), and delete (stop_locally, clear_mock_endpoints, unpublish_from_github). Workflows are chained explicitly (deploy → wait_for_deploy → list_sims). The only notable gap is the absence of update/delete operations for hosted sims themselves — list_sims lists them but nothing can remove or modify an individual hosted sim.
Maintenance
Related MCP Connectors
AI-native mock API server with MCP. Create REST/SOAP mocks from Claude, Cursor, or Windsurf.
Build, validate, and manage API simulations in WireMock Cloud from MCP-compatible AI agents.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Model Context Protocol server for the Apideck Unified API. Connect any MCP-compatible agent framework to 100+ accounting systems, HRIS platforms, file storage providers, and more through one integration. More information https://www.apideck.com/mcp-server
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceThis tool creates a Model Context Protocol (MCP) server that acts as a proxy for any API that has an OpenAPI v3.1 specification. This allows you to use Claude Desktop to easily interact with both local and remote server APIs.620 npm904MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that publishes CLI tools on your machine for discoverability by LLMs10 npm1MIT
- AlicenseAqualityCmaintenanceLocal MCP server that wraps the headless Claude Code CLI as MCP tools, providing stateless access to Claude's coding capabilities through prompt-based interactions. It enables users to execute Claude Code commands with various prompt formats and structured outputs directly from MCP clients.3MIT
- AlicenseNot gradedqualityBmaintenanceA dead simple MCP server for exposing your app functions to AI agents like Claude Desktop.16 npm6MIT