Skip to main content
Glama
Mr-Lucasz

Conversational API Test MCP

by Mr-Lucasz

Conversational API Test MCP

Português

An API test bench your AI agent can drive — and a QA engineer named Vander to drive it.

Most API tools assume a human at the keyboard. This one assumes an agent: you describe your API once in YAML, and the agent runs requests, chains them, and reviews endpoints through a testing heuristic, while the server keeps three promises the agent cannot keep on its own — the request sent is exactly the one declared, the answer costs few tokens, and your credentials never enter the conversation.

It is an MCP server over stdio. It works with Cursor, VS Code + GitHub Copilot, Claude Code, and any other MCP stdio client.

What a session looks like

A real run, recorded on 2026-10-07 against httpbin.org with the definition in examples/httpbin.example.yaml. Only safe methods were sent.

You: /vander .mcp/api/httpbin.yaml whoami

Vander: Vander here. Reviewing GET /bearer on httpbin.org. The plan has 13 automatic checks; 4 of them would send POST, PUT, PATCH or DELETE, so I left those out.

8 passed, 1 failed, 4 skipped.

  • V ✅ OPTIONS → 200. POST / PUT / PATCH / DELETE not sent.

  • A ❌ No credentials → 401, as it should be. But Bearer vander-invalid-token → 200 with "authenticated": true.

  • N ⚪ Nothing to negate: no body, no query parameters, no id in the path.

  • D ✅ Baseline 200, application/json with a JSON body, $.authenticated present.

  • E ✅ No 5xx, and nothing leaked in the 401.

  • R ✅ 196 ms against a 1000 ms budget. One sample, not a load test.

One finding: the endpoint accepts any bearer token. It checks that a token is present, not that it is valid. For an echo service like httpbin that is to be expected; on a real API it is the difference between authentication and a header check.

Not tested: the four state-changing verbs, an expired token, another user's token, rate limiting.

endpoint:
  id: whoami
  method: GET
  path: /bearer
summary:
  pass: 8
  fail: 1
  skipped: 4
  manual: 6
axes[5]:
  - axis: V
    name: Verbs
    checks[5]:
      - id: V-options
        title: OPTIONS does not cause a server error
        result: pass
        expected: status not5xx
        actual: status 200 in 656ms
      - id: V-post
        title: POST (not declared for this path) is rejected
        result: skipped
        note: "sends POST (state-changing); rerun with includeDestructive: true"
      - id: V-put
        title: PUT (not declared for this path) is rejected
        result: skipped
        note: "sends PUT (state-changing); rerun with includeDestructive: true"
      - id: V-patch
        title: PATCH (not declared for this path) is rejected
        result: skipped
        note: "sends PATCH (state-changing); rerun with includeDestructive: true"
      - id: V-delete
        title: DELETE (not declared for this path) is rejected
        result: skipped
        note: "sends DELETE (state-changing); rerun with includeDestructive: true"
  - axis: A
    name: Authorization
    checks[2]:
      - id: A-none
        title: Request without credentials is rejected
        result: pass
        expected: "status in [401, 403]"
        actual: status 401 in 435ms
      - id: A-invalid
        title: Request with an invalid token is rejected
        result: fail
        expected: "status in [401, 403]"
        actual: status 200 in 142ms
        note: "{\"authenticated\":true,\"token\":\"[REDACTED]\"}"
  - axis: D
    name: Data
    checks[3]:
      - id: D-baseline
        title: Request as declared succeeds
        result: pass
        expected: status 2xx
        actual: status 200 in 196ms
      - id: D-content-type
        title: Content-Type matches the body actually returned
        result: pass
        expected: Content-Type consistent with the body
        actual: "content-type \"application/json\", body is JSON"
      - id: D-captures
        title: Every captured JSONPath exists in the response
        result: pass
  - axis: E
    name: Errors
    checks[2]{id,title,result}:
      E-no-5xx,No probe made the server answer 5xx,pass
      E-no-leak,Error bodies do not leak stack traces or database errors,pass
  - axis: R
    name: Responsiveness
    checks[1]{id,title,result,expected,actual}:
      R-duration,Baseline answers within the duration budget,pass,<= 1000ms,196ms

How this was recorded: a script drove the built server over MCP stdio and saved the tool output shown above, unedited. Vander's reply was then written by Claude from that output, following the vander prompt. Every number in the reply comes from the output; timings will differ on your machine.

Related MCP server: fetchsandbox-mcp

Three bets

1. The agent chooses; the server executes. The agent picks a requestId, and the server builds and sends the request from the YAML. There is no hand-written curl in the chat to get subtly wrong, and the same call gives the same request tomorrow.

2. Tokens are a budget. responseDetail, jsonPathSelect and maxBodyChars decide how much of a response comes back; execute_api_flow runs a whole scenario in one call; results are encoded as TOON rather than JSON.

3. A heuristic is worth more than a personality. Vander is a persona, but what he does is fixed: a checklist built by code and verdicts computed from status codes and timings. Research on personas in prompts found they do not, by themselves, make a model more accurate (sources), so the personality is there for the conversation and the rigour lives in run_vander_checks.

Quick start

Requires Node.js ≥ 20.

npm ci
npm run build

Register the server in your MCP client, pointing at dist/index.js:

{
  "servers": {
    "conversational-api-test-mcp": {
      "type": "stdio",
      "command": "node",
      "args": ["/absolute/path/to/conversational-api-test-mcp/dist/index.js"]
    }
  }
}

In the project you want to test, create .mcp/api/ and add a definition (see examples/default.example.yaml):

version: "1"
service: weather
base_url: "{{BASE_URL}}"
flows:
  smoke:
    steps:
      - get_token
      - requestId: forecast
        assert:
          jsonPathExists: $.days
endpoints:
  - id: get_token
    method: POST
    path: /oauth/token
    form:
      grant_type: client_credentials
      client_id: "{{CLIENT_ID}}"
      client_secret: "{{CLIENT_SECRET}}"
    capture:
      TOKEN: $.access_token
  - id: forecast
    method: GET
    path: /forecast
    params:
      city: lisbon
    auth: Bearer {{TOKEN}}
    auth_dependency: get_token

Put secrets in .env.mcp.local at the workspace root and never commit it:

BASE_URL=https://api.example.com/v1
CLIENT_ID=...
CLIENT_SECRET=...
STAGING_BASE_URL=https://staging.example.com/v1

Then ask the agent explicitly, e.g. "run the smoke flow from .mcp/api/weather.yaml through the MCP".

Tools

Tool

Purpose

list_api_definitions

List YAMLs under .mcp/api/ (glob, pagination, sort, field projection).

summarize_api_definition

Cheap triage: endpoint ids, variable keys, flow names.

read_api_definition

Full, validated definition.

execute_api_request

Run one endpoint.

execute_api_flow

Run several endpoints in one call — inline steps or a flowName declared in the YAML.

dry_run_request

Show the request that would be sent, without sending it.

assert_response

Assert on the last response (status, jsonPathExists).

set_environment

Select the active environment (CURRENT_ENV).

set_environment_variable / get_environment_variable

Session variables.

explain_request_context

Which variable keys are available from where.

upsert_canonical_api_definition

Create / append definitions (dry-run by default).

reorganize_mcp_api_definitions

Merge many YAMLs into fewer (plan, then apply).

plan_vander_checks / run_vander_checks

VANDER review of one endpoint: checklist, then automatic execution grouped by axis.

discover_legacy_api_sources / convert_legacy_to_canonical

Import Postman, OpenAPI or Insomnia files.

Recommended order: list_api_definitions → summarize_api_definition → set_environment (if needed) → execute_api_request or execute_api_flow.

Vander

Vander is a persona shipped as an MCP prompt: a senior API QA engineer you talk to instead of issuing tool calls yourself. In clients that expose MCP prompts it shows up as a command (for example /vander), optionally with a definition path and an endpoint id. He reviews an endpoint through the VANDER heuristic:

Axis

Question

Checked automatically

Verbs

What does the path do with methods it does not declare?

Undeclared methods are rejected (405/404/501); OPTIONS does not error.

Authorization

Who can call it?

No credentials and an invalid token are rejected (401/403).

Negative

What happens with bad input?

Each top-level body field removed in turn is rejected (4xx).

Data

Does the response say what it should?

Baseline succeeds, Content-Type matches the body, captured JSONPaths exist.

Errors

Does it fail well?

Malformed JSON is a client error; no probe causes a 5xx or leaks a stack trace.

Responsiveness

Is it fast enough?

Baseline answers within maxDurationMs (default 1000).

Each axis also carries manual ideas (other users' tokens, boundary values, idempotency, rate limiting…) that Vander explores with the regular tools.

The two tools work without the prompt too:

  • plan_vander_checks returns the checklist for a requestId and sends nothing.

  • run_vander_checks executes the automatic checks and returns pass / fail / skipped per axis. Probes never capture variables. Anything that sends POST, PUT, PATCH or DELETE is skipped unless includeDestructive: true, so reviewing a write endpoint is an explicit decision.

VANDER builds on Stuart Ashman's VADER heuristic and adds an explicit Negative axis — see where the ideas come from.

Definitions

Variables and environments

{{KEY}} is resolved from, in increasing priority: YAML variables → .env.mcp.local → session variables (captures and set_environment_variable).

  • After set_environment (e.g. STAGING), {{BASE_URL}} resolves STAGING_BASE_URL first and falls back to BASE_URL.

  • Unresolved placeholders are left as literal {{KEY}} so the problem is visible.

  • Macros: {{$uuid}}, {{$timestamp}}, {{$date}}, {{$date:YYYY-MM-DD HH:mm:ss}}.

  • {{env.NAME}} reads the server's process.env, but only for names listed in MCP_API_ENV_PASSTHROUGH.

Requests

  • base_url may carry a path prefix (https://host/api/v1); path is appended to it. An absolute path overrides base_url.

  • params is sent as the query string; headers, body (JSON) and form (URL-encoded) are interpolated.

  • capture maps a session variable name to a JSONPath on the response; assert checks status and/or jsonPathExists.

Auth

  • auth: Bearer {{TOKEN}} sets the Authorization header.

  • auth_dependency: <endpoint id> runs that endpoint first when its captured variables are missing, and again after a 401 (disable with auth_retry_on_401: false).

  • digest_auth: { username, password } performs HTTP Digest.

Flows

Each step is an endpoint id or an object:

flows:
  create_and_wait:
    steps:
      - create_job
      - requestId: get_job
        poll: { untilJsonPath: "$.finishedAt", maxAttempts: 10, delayMs: 2000 }
        retry: { max: 3, delayMs: 500 }
        acceptStatus: [200]
        assert: { status: 200, jsonPathExists: "$.result" }
        optional: false

Keeping output small

  • responseDetail: "minimal" returns status, duration and captures only; "summary" (default) adds a body preview capped at 8,000 chars; "full" raises the cap to 50,000.

  • jsonPathSelect projects the body before it is serialized.

  • execute_api_flow replaces N tool calls with one.

  • Parsed definitions and .env.mcp.local are cached in memory and invalidated by file mtime.

Security model

The server runs locally with your privileges and is driven by an LLM, so treat API responses and definitions as untrusted input.

  • Redaction. Auth headers, values under credential-looking JSON keys (password, token, secret, api_key, …), JWT-looking strings, credential query parameters and captured credentials are returned as [REDACTED]. The real values stay in the session and are still used by later requests. Redaction is name- and shape-based, so it is a safety net, not a guarantee.

  • No implicit environment access. process.env is not part of the interpolation context; only names you list in MCP_API_ENV_PASSTHROUGH are reachable, via {{env.NAME}}.

  • Responses are data. Values captured from a response are inserted literally and never re-expanded as templates.

  • Writes are confined to .mcp/api/ and default to dry-run.

  • Optional host allowlist. MCP_API_ALLOWED_HOSTS restricts requests — including every redirect hop — to the listed hostnames. It is off by default and is not a complete SSRF defence (no private-range or DNS-rebinding checks).

Environment variable

Effect

MCP_WORKSPACE_ROOT

Default workspaceRoot when a tool call omits it.

MCP_API_ALLOWED_HOSTS

Comma-separated hostnames requests may target. Unset = any host.

MCP_API_ENV_PASSTHROUGH

Comma-separated env var names (trailing * = prefix) exposed to {{env.NAME}}. Unset = none.

MCP_API_MAX_RESPONSE_BYTES

Response size cap. Default 10 MB.

MCP_API_REVEAL_SECRETS

true disables redaction. Local debugging only.

Where the ideas come from

Nothing here was invented from scratch. This is what the project borrows, from whom, and what it changes.

Idea in this project

Source

What was taken, and what was changed

The VANDER axes

Stuart Ashman, VADER – a REST API test heuristic (QA Matters, 2016)

The five original axes — Verbs, Authorization, Data, Errors, Responsiveness — are his. This project adds Negative as a sixth and turns part of each axis into executable checks.

The Negative axis

Gwen Diagram & Ash Winter's BINMEN (Boundary, Invalid entries, NULL, Method, Empty, Negative), as listed in the Ministry of Testing Test Heuristics Cheat Sheet

The framing of negative input as its own concern. Only "missing field" is automated today; the rest are offered as manual ideas.

Splitting input from output concerns

Amber Race's POISED (Parameters, Output, Interop, Security, Errors, Data), same cheat sheet

Informs the manual ideas under Data and Negative.

Which probes are "destructive"

RFC 9110, HTTP Semantics §9.2.1 (safe methods)

GET, HEAD and OPTIONS are sent freely; every other method needs includeDestructive.

Expected statuses

RFC 9110 §15.5.6 (405), §15.6.2 (501), §15.5.2 (401), §15.5.4 (403)

An undeclared method should get 405 or 501 (404 is also accepted); missing or bad credentials should get 401 or 403.

What the Authorization axis looks for

OWASP API Security Top 10 – 2023, API2 Broken Authentication

Automated: no credentials, invalid token. API1 and API5 (object- and function-level authorization) need a second identity, so they stay manual.

Host allowlist, response size cap

OWASP API7 (Server Side Request Forgery) and API4 (Unrestricted Resource Consumption)

Applied to this server itself, since it makes requests on an agent's behalf.

Vander as an MCP prompt

MCP specification, Prompts

Prompts are "user-controlled": the user picks them explicitly, typically as a slash command. That is why the persona is opt-in rather than baked into every conversation.

Persona for tone, code for verdicts

Zheng et al., When "A Helpful Assistant" Is Not Really Helpful (Findings of EMNLP 2024)

They report that adding a persona to the system prompt did not improve accuracy over no persona. So nothing Vander asserts depends on the persona.

TOON output

TOON specification

Used as-is for tool results. Its authors report the largest savings on uniform arrays and little or none on deeply nested data; this project has not benchmarked it independently.

VANDER is this project's name for the extended heuristic and has no affiliation with the authors above.

Troubleshooting

Symptom

Check

Empty definition list

Is workspaceRoot absolute and correct? Does .mcp/api/ contain .yaml files?

401

set_environment, the endpoint's auth_dependency, captures, .env.mcp.local.

{{KEY}} sent literally

Key name, environment prefix (STAGING_*), or — for shell variables — MCP_API_ENV_PASSTHROUGH.

A value shows as [REDACTED]

Expected for credentials. It is still used in requests.

Development

Command

Description

npm run build

Compile TypeScript to dist/.

npm run dev

Run from source with tsx.

npm test

Run the test suite (vitest).

npm run lint

ESLint.

License

MIT

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables AI agents to dynamically discover and interact with APIs through Swagger/OpenAPI specifications and Postman collections using a strategic four-tool approach. It streamlines API integration by providing universal tools for endpoint discovery, detailed request information, and authenticated execution.
    1
    -
  • A
    license
    A
    quality
    A
    maintenance
    A deterministic eval engine for coding agents. Your agent claims it fixed the bug—this checks it. The same scenario runs twice against a sandbox of the services your code calls, with failures injected on purpose. It must fail on the old code and pass on the new. The verdict is an exit code, not a model's opinion. Every run leaves a receipt.
    8
    16
    433 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to inspect, test, validate, compare, and debug HTTP APIs through a structured toolkit, with features such as secret redaction, response contract validation, regression comparison, health checks, and incident investigation workflows.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to call real API endpoints and verify responses against OpenAPI specs, supporting contract testing, auth presets, spec diffing, and health checks.
    8
    52 npm
    MIT