api-test-mcp
# api-test-mcp
<img src="assets/banner.svg" alt="api-test-mcp — call the real API, check it against what the spec promises" width="100%">
[](https://github.com/thisis-najeeb/api-test-mcp/actions/workflows/ci.yml)
[](https://www.npmjs.com/package/api-test-mcp)
[](https://www.npmjs.com/package/api-test-mcp)
[](LICENSE)
An MCP server that gives Claude Code, Cursor, Windsurf, or any MCP-compatible AI agent the ability to **actually call your API and check the response against what your OpenAPI spec promises** — not just read the docs and guess.
No API keys, no config, no cost. Works with any OpenAPI/Swagger 3.x spec (URL or local file).
## Why this exists
AI agents are great at reading an OpenAPI spec and writing code against it — but they're guessing about whether the real API actually behaves the way the spec says. This gives an agent (or you, in a normal chat) a way to find out for real: call the live endpoint, and check whether the response actually matches the documented schema.
## Install
```bash
git clone <this repo>
cd api-test-mcp
npm install
```
Add it to your MCP client config, e.g. for Claude Code:
```bash
claude mcp add api-test -- node /absolute/path/to/api-test-mcp/src/index.js
```
Or in `claude_desktop_config.json` / Cursor's MCP settings:
```json
{
"mcpServers": {
"api-test": {
"command": "node",
"args": ["/absolute/path/to/api-test-mcp/src/index.js"]
}
}
}
```
## Tools
| Tool | What it does |
|---|---|
| `load_api_spec` | Load and dereference an OpenAPI/Swagger spec from a URL or local path. Returns the API title, servers, and every documented endpoint. Call this first. |
| `list_endpoints` | List every endpoint currently loaded. |
| `call_endpoint` | Make a real HTTP call to a documented endpoint. Returns the real status, headers, and body. |
| `validate_response` | Check a response body against the JSON schema documented for a given method + path + status. |
| `test_endpoint` | `call_endpoint` + `validate_response` in one step. The main tool — "does this endpoint actually work as documented?" |
| `run_all_tests` | Best-effort contract-test pass across every GET endpoint that needs no required parameters. Pass `includeMutating: true` to also auto-generate example params/bodies from the schema and attempt POST/PUT/PATCH (off by default — it can write real data). Endpoints still needing manual input are listed as skipped, with the reason. |
| `check_health` | One-shot ping across a set of endpoints (or every parameter-free GET in the loaded spec): reports reachability and latency. Handy before a demo or as a CI step. |
| `diff_api_specs` | Compare two versions of a spec (e.g. an old tag vs. `main`) and flag likely-breaking changes — removed endpoints, newly-required fields, type changes, removed enum values — versus safe additive changes. |
All of the above accept an optional `auth` preset (`bearer`, `apiKey` in a header or query param, or `basic`) so authenticated APIs aren't limited to hand-building raw headers, and an optional `timeoutMs`.
## Example (what an agent conversation looks like)
> **You:** Load my API spec at `https://api.example.com/openapi.json` and check whether `/users/{id}` actually returns what it documents.
>
> **Agent:** *(calls `load_api_spec`, then `test_endpoint` with a real user id)* → "Called it — got a 200, but the response is missing the `created_at` field your spec marks as required, and `role` is documented as an enum of 3 values but the API returned `"superadmin"`, which isn't one of them."
For an authenticated API:
> **You:** Run a full contract-test pass against my staging API using this bearer token, and include the write endpoints.
>
> **Agent:** *(calls `run_all_tests` with `{ auth: { type: "bearer", token: "..." }, includeMutating: true }`)* → "12 passed, 2 failed, 3 skipped. `POST /orders` failed schema validation — `total_cents` came back as a string, not the integer your spec documents."
## Tested against real live traffic
`npm test` runs three real, unmocked checks, no canned fixtures pretending to be a server:
- `test/smoke-test.js` — loads a spec, makes real HTTPS calls to a live public API, validates the real response, and deliberately feeds in a broken response to confirm validation actually catches mismatches (not just a happy-path check).
- `test/new-features-test.js` — auth presets applied to a real outgoing request URL, a real network timeout/abort, a real health-check call, and deterministic offline tests for the spec-diff logic.
- `test/mcp-protocol-test.js` — spawns the actual MCP server as a subprocess and talks to it over the real MCP protocol, the same way Claude Code or Cursor would.
CI runs the full suite on every push/PR against Node 18, 20, and 22.
## Roadmap
v1.0 shipped contract testing, auth presets, auto-generated example data for mutating endpoints, spec diffing, and health checks. Ideas for what's next:
- YAML output mode / a small CLI wrapper for non-MCP use
- Configurable retry/backoff for flaky endpoints in `run_all_tests` and `check_health`
- Pattern-aware example generation (respect JSON Schema `pattern` instead of a placeholder string)
- Persisted health-check history (currently one-shot only)
Contributions welcome — see [CONTRIBUTING.md](CONTRIBUTING.md). See an endpoint type or spec quirk this doesn't handle well? Open an issue.
## License
MIT
TDQS
Scored across 8 tools
Tools generally target distinct phases (load, list, call, validate, batch-test, health-check, diff), but call_endpoint/test_endpoint and run_all_tests/check_health both invoke live endpoints and could be confused if the goal is simply 'hit the API'. Descriptions do enough to resolve most ambiguity, so not a 5.
All tools use lowercase snake_case imperative verbs (load_, list_, call_, validate_, test_, run_, check_, diff_), and each name clearly signals the action. The pattern is highly consistent and predictable.
8 tools is well-scoped for an API testing server: loading, exploring, calling, validating, running tests, health-checking, and diffing specs. No redundant or extraneous tools are present.
The set covers the full workflow from loading an API spec to listing endpoints, making raw calls, validating responses, batch-running endpoint tests, health-checking, and comparing spec versions. Additional tools like auth management or report generation would be optional rather than obvious gaps.