Skip to main content
Glama

Backwork MCP Server

npm

Official Model Context Protocol (MCP) server for the Backwork API. It gives AI assistants controlled access to Medicare and payer medical policies, medical code intelligence, prior authorization checks, claim validation, compliance review, drug formulary evidence, and webhook operations.

There are two ways to connect:

Hosted remote

Local stdio

Endpoint

https://backworkhealth.com/mcp (Streamable HTTP)

npx -y @backwork/mcp

Auth

OAuth in the browser, no key to copy

BACKWORK_API_KEY=bwk_live_...

Use when

Your client supports remote MCP with OAuth

Your client only runs local commands, or you want the server on your machine

The hosted endpoint only accepts OAuth. Do not send a Backwork API key as a bearer token to https://backworkhealth.com/mcp. The OAuth grant is read-only (backwork:mcp read), so the hosted server offers only read actions: no webhook management and no compliance acknowledgements. Use local stdio with a write-scoped live key for those.

Calls draw on your organization's request credits, like any other /api/v1 call. A bwk_test_ key works only on the sandbox (https://backworkhealth.com/api/sandbox/v1), which covers policy search, code lookup, prior-auth check and coverage evaluation; to use it, set BACKWORK_API_BASE to that URL.

Claude Code

claude mcp add backwork --transport http https://backworkhealth.com/mcp

Add --scope user to make it available in every project. Then run claude, open /mcp, select backwork, and finish the browser login and Backwork consent screen. Check it with claude mcp get backwork.

If OAuth discovery needs to be pinned explicitly, add the same server as JSON:

claude mcp add-json backwork '{
  "type": "http",
  "url": "https://backworkhealth.com/mcp",
  "oauth": {
    "scopes": "backwork:mcp read"
  }
}'

Local stdio:

claude mcp add backwork -e BACKWORK_API_KEY=bwk_live_YOUR_API_KEY -- npx -y @backwork/mcp

Related MCP server: mymedi-ai-mcp-server

Claude Desktop

Hosted remote: open Settings > Connectors > Add custom connector, name it Backwork, and enter https://backworkhealth.com/mcp. Claude Desktop runs the OAuth login when you connect.

Local stdio: add this to claude_desktop_config.json (Settings > Developer > Edit Config) and restart Claude Desktop:

{
  "mcpServers": {
    "backwork": {
      "command": "npx",
      "args": ["-y", "@backwork/mcp"],
      "env": {
        "BACKWORK_API_KEY": "bwk_live_YOUR_API_KEY"
      }
    }
  }
}

Cursor

Add to ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project). Cursor opens the OAuth login the first time it connects:

{
  "mcpServers": {
    "backwork": {
      "url": "https://backworkhealth.com/mcp"
    }
  }
}

Local stdio:

{
  "mcpServers": {
    "backwork": {
      "command": "npx",
      "args": ["-y", "@backwork/mcp"],
      "env": {
        "BACKWORK_API_KEY": "bwk_live_YOUR_API_KEY"
      }
    }
  }
}

VS Code

code --add-mcp '{"name":"backwork","type":"http","url":"https://backworkhealth.com/mcp"}'

Or add it to .vscode/mcp.json in a workspace. VS Code asks you to sign in when the server starts:

{
  "servers": {
    "backwork": {
      "type": "http",
      "url": "https://backworkhealth.com/mcp"
    }
  }
}

Local stdio, with the key prompted for once and stored by VS Code:

{
  "inputs": [
    {
      "type": "promptString",
      "id": "backwork-api-key",
      "description": "Backwork API key",
      "password": true
    }
  ],
  "servers": {
    "backwork": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@backwork/mcp"],
      "env": {
        "BACKWORK_API_KEY": "${input:backwork-api-key}"
      }
    }
  }
}

Codex

codex mcp add backwork --env BACKWORK_API_KEY=bwk_live_YOUR_API_KEY -- npx -y @backwork/mcp

Use in ChatGPT

ChatGPT connects to the hosted server as a custom MCP connection in developer mode. Three tools return a small card that ChatGPT renders inline above its answer.

  1. In ChatGPT, open Settings → Security and login and turn on Developer mode. Your plan or workspace admin may need to allow it.

  2. Go to chatgpt.com/plugins and select the + button.

  3. Name it Backwork, add a short description, and under Connection enter https://backworkhealth.com/mcp. Choose OAuth for authentication.

  4. Select Create. ChatGPT opens Backwork's sign-in page; approve the read-only backwork:mcp read grant.

  5. Check the discovered tools, then start a new chat, add Backwork from the tools menu, and ask something like "Is CPT 76942 covered in Texas, and does it need prior auth?"

After a server update, open the connection at chatgpt.com/plugins and select Refresh so ChatGPT picks up new tool and component metadata.

Tool

Card

What it shows

backwork_coverage_lookup

Coverage result

Each policy that lists the codes: payer (CMS for Medicare policies), title and number, effective date, a row per code with its disposition badge and source label, and an Open policy link to the policy's Backwork page (Medicare LCDs, Articles, and NCDs) or, for other policies, its source document

backwork_prior_auth_research

Prior-auth checklist

The determination and confidence, codes that require prior auth, documentation to gather, known gaps, and numbered citations. A started research task shows as pending until you ask again

backwork_policy_research

Policy research

Per action: compare shows the codes side by side across Medicare contractors (MACs), one column per jurisdiction, with each cell's disposition and policy, and a coverage count per column; search lists the matching policies with payer, number, effective date, a short summary and a link; get shows one policy's summary, a criteria excerpt per section, and its codes; criteria lists the matching criteria excerpts with their policies; changes lists recent changes; jurisdictions lists each MAC and its states

ChatGPT loads a tool's card for every call of that tool. A result with nothing to show, such as a search with no matches or an error, collapses the card to zero height and asks ChatGPT to close it, so no empty card is left in the conversation.

A code that a policy lists only because its title names the drug is labelled Inferred from policy title, in the card and in the text. Confirm it against the document before relying on it.

The cards follow ChatGPT's light or dark theme, load nothing from the network (their CSP allows no domains), and open links through the host. They are MCP Apps resources (text/html;profile=mcp-app), so other MCP Apps hosts can render them too. Clients without UI support ignore the _meta keys that link tools to cards and get the same markdown text as before; structuredContent.widget carries the card's data.

Other MCP Clients

Clients that run local commands can use the stdio config shown for Claude Desktop. Clients that support remote URLs with OAuth can use https://backworkhealth.com/mcp directly.

For clients that support only remote URLs with static headers, deploy a private self-hosted server in API-key or dual-auth mode (see Self-Hosting) and send the key as a bearer header:

{
  "mcpServers": {
    "backwork": {
      "url": "https://your-private-mcp.example.com/mcp",
      "headers": {
        "Authorization": "Bearer bwk_live_YOUR_API_KEY"
      }
    }
  }
}

Self-Hosting

Run a Streamable HTTP server:

git clone https://github.com/tylergibbs1/backwork-mcp.git
cd backwork-mcp
npm install
npm run build
npm run start:http

Defaults:

Setting

Default

Override

Transport

stdio

--http or BACKWORK_MCP_TRANSPORT=http

Host

127.0.0.1

--host or BACKWORK_MCP_HOST

Port

3000

--port or BACKWORK_MCP_PORT or PORT

MCP path

/mcp

--path or BACKWORK_MCP_PATH

Allowed hosts

loopback/private hosts, VERCEL_URL, or configured public host

BACKWORK_MCP_ALLOWED_HOSTS or BACKWORK_MCP_PUBLIC_HOST

HTTP mode requires Authorization: Bearer per request. By default this bearer is a Backwork API key. For hosted remote MCP deployments, enable OAuth protected-resource discovery so Claude-compatible clients can authenticate users through your authorization server:

BACKWORK_MCP_AUTH_MODE=oauth \
BACKWORK_MCP_OAUTH_AUTHORIZATION_SERVERS=https://backworkhealth.com \
BACKWORK_MCP_OAUTH_SCOPES="backwork:mcp read" \
npm run start:http

The server publishes OAuth Protected Resource Metadata at /.well-known/oauth-protected-resource and includes that URL in WWW-Authenticate challenges. If your Backwork API accepts OAuth access tokens directly, no extra mapping is needed; the MCP server forwards the OAuth bearer downstream. If your authorization server exposes a Backwork API key in token introspection, set BACKWORK_MCP_OAUTH_INTROSPECTION_URL and BACKWORK_MCP_OAUTH_API_KEY_CLAIM to validate the access token and map it to the downstream Backwork credential.

For a private single-tenant deployment where the server environment supplies the key, set:

BACKWORK_MCP_ALLOW_ENV_KEY=true BACKWORK_API_KEY=bwk_live_YOUR_API_KEY npm run start:http

Only use BACKWORK_MCP_ALLOW_ENV_KEY=true on loopback or private-network deployments protected by network access control. Public deployments should require a bearer token per request, set BACKWORK_MCP_ALLOWED_HOSTS/BACKWORK_MCP_PUBLIC_HOST, and set BACKWORK_MCP_ALLOWED_ORIGINS only to exact browser origins that may connect.

Vercel Hosting

This repo can deploy as an API-only Vercel project. The production project uses:

BACKWORK_MCP_AUTH_MODE=oauth
BACKWORK_MCP_PUBLIC_HOST=backworkhealth.com
BACKWORK_MCP_PUBLIC_URL=https://backworkhealth.com
BACKWORK_MCP_ALLOWED_HOSTS=backworkhealth.com,mcp.backworkhealth.com,backwork-mcp.vercel.app
BACKWORK_MCP_OAUTH_AUTHORIZATION_SERVERS=https://backworkhealth.com
BACKWORK_MCP_OAUTH_RESOURCE=https://backworkhealth.com/mcp
BACKWORK_MCP_OAUTH_SCOPES="backwork:mcp read"
BACKWORK_MCP_OAUTH_REQUIRED_SCOPES=backwork:mcp
BACKWORK_MCP_OAUTH_INTROSPECTION_URL=https://backworkhealth.com/api/oauth/introspect
BACKWORK_MCP_OAUTH_EXPECTED_AUDIENCE=https://backworkhealth.com/mcp

The Vercel functions expose:

Path

Purpose

/mcp

Streamable HTTP MCP endpoint

/health

Lightweight MCP server health check

/.well-known/oauth-protected-resource

OAuth protected-resource metadata when OAuth is configured

/

Basic endpoint metadata

The Backwork web app that issues OAuth tokens must also be configured:

BACKWORK_OAUTH_ISSUER=https://backworkhealth.com
BACKWORK_OAUTH_SIGNING_SECRET=<generate with: openssl rand -base64 48>
BACKWORK_MCP_RESOURCE=https://backworkhealth.com/mcp

Production OAuth discovery fails closed unless BACKWORK_OAUTH_SIGNING_SECRET is at least 32 characters and Redis or Vercel KV is configured for one-time consent and authorization-code storage.

Health check:

curl http://localhost:3000/health

Local Development

npm install
npm run build
BACKWORK_API_KEY=bwk_live_YOUR_API_KEY npm start

Useful commands:

npm run start:http
node build/src/index.js --help

Requires Node.js 18.14.1 or newer.

Available Tools

Tool names use the backwork_ prefix for discoverability when this server is installed alongside other MCP servers. The default surface is intentionally workflow-level rather than a 1:1 API wrapper, so agents see fewer choices and common tasks require fewer tool calls.

All tools include title, description, inputSchema, outputSchema, and MCP annotations. Each tool states its own annotations; the server refuses to start if a tool marked read-only offers a write-scope operation, or one marked non-destructive offers a DELETE. Successful calls return readable text plus structuredContent with message, and when available, Backwork API data and meta.

data and meta are projections: they carry only the response fields src/api-operations.ts lists for the operation (reads and metaReads). Request IDs, response timestamps, record timestamps, research-job cost and polling URLs, idempotency keys, and model names are dropped. Policy effective and last-reviewed dates, source fetch times, and source audit sample times are kept. The server logs the request ID of a failed API call.

Tool-level failures return isError: true. Billing and rate-limit failures are reported in plain terms and never repeat the API's hint, plan names, or pricing links: a feature outside the organization's plan, no remaining request credits, or a rate limit with when to retry. test/output-guard.test.mjs runs every tool action against fixture responses and those errors, and fails on upsell wording, trace IDs, timestamps, cost, or model fields.

Sources and currency

When the Backwork API response cites policy documents, structuredContent.provenance carries what an agent needs to cite them, and the text output ends with a short --- Sources --- block:

Field

Meaning

source_urls

Distinct source document URLs

authorities

Issuing authorities, for example CMS or a payer name

retrieved_at

Oldest fetch time when every cited source has a known time; otherwise null

as_of

Latest effective date among the cited sources

sources[]

policy_id, source_url, authority, retrieved_at, as_of per source

Values come from the API response. source_check.source_url and source_check.last_fetched_at take precedence over legacy source metadata. An explicit unknown fetch time stays null; legacy retrieved_at, last_verified_at, or crawled_at are fallback values only when the source-check block is absent. Known fetch times remain available in sources[] when another cited source has an unknown time. The shape matches the provenance block on Backwork agent tool results.

Policy evidence also retains the API's source_check: source URL, last fetch time, fetched-content SHA-256, and source_accuracy audit samples. Each audit reports the field measured, sample date and size, share of sampled records that matched, 95% Wilson interval, and method. These samples measure records from a source; they do not measure certainty for the individual policy or a coverage decision. Missing audits stay null. Cards show fetch dates and expandable source sample audits.

Policies with explicit applicability also retain the optional applicability_scope, applicability_markets, applicability_evidence and applicability_note fields. Shared-market documents and indexed singleton documents are listed in the publisher's state indexes; that discovery evidence does not establish the member's plan or product coverage. A document_scoped policy with no state, line of business or indexed markets has unknown applicability. The cards and text show the API's note, and cards provide publisher index links and source-exact quotes with page numbers. Legacy responses without these fields keep their existing shape.

Applicability evidence is limited to a document SHA-256, up to 16 HTTPS publisher index listings and up to 16 statements of the three supported kinds (non_medicare_disclaimer, commercial_policy_header, member_type_branch). Quotes are bounded to 4,000 characters, index URLs to 2,048 characters, and pages must be positive integers. Unknown fields are dropped; malformed evidence becomes null. Market codes and notes are bounded as well. The MCP preserves API confidence and manual-review outputs; index listings do not raise confidence or resolve authorization.

Code matches retain grounding separately from their extraction source:

grounding

Meaning

grounded

Found in Backwork's retained source text

not_grounded

Read by the document pipeline, but not found in retained source text

no_source_text

No retained text was available to check

not_checked

Grounding was not checked, including codes inferred from a policy title

Text and cards label these checks without treating a grounded code as a coverage guarantee. An older response that omits grounding remains unknown, and an unrecognized future value is preserved and labelled.

Production availability

Which actions a server offers depends on two things:

  • Availability. The Backwork OpenAPI document can mark an operation x-backwork-availability: unavailable-in-production. Against the production API (https://backworkhealth.com) the server withholds those actions. None is marked today: production serves every /api/v1 operation these tools call to an organization's live key.

  • Access. Operations that need write scope (x-backwork-required-scopes) are withheld from a read-only OAuth grant. On the hosted server this hides backwork_webhook_management and the acknowledge and bulk_acknowledge actions of backwork_compliance_review. A read-only connection also gets no backwork_system_health diagnostics, no idempotency_key input on backwork_claim_validation, and a backwork_compliance_review with no acknowledgment inputs (diff_id, diff_ids, notes) and no diff IDs in its results. A Backwork API key (stdio, or HTTP with a bwk_ bearer) is offered every tool and action, and the API enforces the key's own scopes. An OAuth grant counts as read-only unless its scopes include write.

A tool with no offered action is hidden. A partly offered tool drops the withheld actions from its action input; actions withheld as unavailable in production are also named in its description. A server pointed at another Backwork deployment with BACKWORK_API_BASE ignores availability markers, and BACKWORK_MCP_EXPOSE_UNAVAILABLE_TOOLS=true does the same against production. Neither lifts the write-scope rule.

Primary tool

Purpose

backwork_coverage_lookup

Look up procedure codes and combine code details, policy evidence, prior authorization, claim risk, jurisdiction comparison, and spending evidence. With payer, code details list only that payer's policies and say how many other payers' policies were left out

backwork_policy_research

Search policies, fetch one policy, search extracted criteria, review policy changes, map MAC jurisdictions, or compare how MACs cover the same codes

backwork_claim_validation

Validate claim coverage, documentation requirements, denial risk, and optional policy-specific criteria

backwork_prior_auth_research

Check prior authorization from Backwork's policies (Medicare, or a named payer's), or start and poll a background job that searches public payer websites

backwork_drug_formulary_research

Search commercial pharmacy-benefit evidence from CVS Caremark, Express Scripts, and UnitedHealthcare / Optum Rx

backwork_compliance_review

Review compliance stats and list unreviewed policy changes; with an API key, also acknowledge changes

backwork_webhook_management

List, create, update, delete, or test webhook endpoints. Backwork sends one event, compliance.acknowledged; policy-change webhooks are not sent. Needs a write-scoped API key

backwork_system_health

Check Backwork API health and dependency status. API-key connections only

Response Format

Every tool accepts:

{
  "response_format": "markdown"
}

Use "markdown" for readable output or "json" to make the text content mirror the returned structuredContent.

Example Prompts

Is CPT 76942 covered in Texas, and does it require prior authorization?
Compare coverage for J0585 across JM and JH.
Validate denial risk for 99213 with diagnosis E11.9 for Medicare in Texas.
Search formulary evidence for Ozempic across commercial PBMs.

Testing and Evaluations

Run the build and MCP metadata smoke test:

npm test

The smoke test starts the built stdio server with a dummy key, verifies the 8 workflow tools, checks titles, schemas, annotations, output schemas, response_format, and verifies local validation failures are reported with isError: true. The HTTP smoke also completes authenticated MCP initialization, lists the read-grant tools, and calls a tool with invalid arguments. The test/ suite covers production availability, the hosted server's OAuth-only, read-only tool set, provenance, description quality and a lexical tool-selection check (no model calls), the registry manifest, and the ChatGPT cards: their templates in the tool list, the resources and mime type, each card's structuredContent against its schema, markdown for clients without UI, and a render of each card page in a stub browser for every action of every tool with a card, which must show a card or collapse to zero height.

Before publishing, run npm run verify:package. It installs the packed tarball in a temporary consumer without repository overrides or a lockfile, then runs both transports on the current runtime and Node 18.20.8. CI runs the same check. Additional exact Node versions can be passed after --.

API contract

Every Backwork endpoint this server calls is listed in src/api-operations.ts, with the request fields it sends and the response fields it reads. Tools can only call the API through that catalog. Each entry also mirrors the operation's availability marker and required scope. npm test checks the catalog against the vendored openapi/backwork-openapi.json, and asks the server's own exposure rule which tool actions it would offer on production, for a read-only OAuth grant and for an API key. An action offered for an operation that production does not serve, marks unavailable, or (for the OAuth grant) guards with write scope fails.

CI runs npm run contract:live, which runs the same checks against https://backworkhealth.com/openapi.json and fails when the vendored copy differs from it anywhere except descriptions, summaries and examples.

When the Backwork API changes, refresh the vendored copy and review the diff:

npm run openapi:update
npm run contract:check

The evals/ directory includes a tool-discoverability evaluation and a read-only data evaluation built from fixed source-backed policy/code records. Refresh the read-only answers intentionally when Backwork source data is updated.

SDK migration follow-up

Release 2.1.4 bundles the v1 SDK 1.32.1, Zod, and their production dependencies from the committed lockfile. This carries the patched Hono Node adapter 1.19.15 into consumer installs and preserves Node 18 support (minimum 18.14.1). npm only honors overrides in the consumer's root package, so an override in this package cannot provide that guarantee. Use npm ci to reproduce the release tree and rerun package verification after updating the lockfile. The official v2 migration guide requires a separate compatibility change:

  1. Raise the supported Node.js runtime from 18 to at least 20, including the hosted deployment and release workflow.

  2. Replace monolithic SDK imports with @modelcontextprotocol/server, @modelcontextprotocol/node for Node HTTP transport, and @modelcontextprotocol/client for clients and tests; use @modelcontextprotocol/core for public protocol schemas where needed.

  3. Upgrade the declared Zod range to at least 4.2 and wrap tool input/output schemas as Standard Schema objects, preserving field descriptions and structured-output validation.

  4. Verify tool discovery, stdio, stateless HTTP lifecycle, OAuth discovery and scopes, widgets, output minimization, and the live OpenAPI contract before publishing that migration.

MCP Registry

server.json describes this server for the official MCP registry as io.github.tylergibbs1/backwork-mcp: the hosted Streamable HTTP remote at https://backworkhealth.com/mcp (OAuth, discovered from the protected-resource metadata) and the @backwork/mcp npm package over stdio. package.json carries the matching mcpName that the registry uses to verify npm ownership. npm test validates server.json against the registry schema and checks that its versions match package.json.

Release

Releases publish @backwork/mcp to npm with Trusted Publishing (GitHub OIDC, no npm token) and then publish server.json to the MCP registry.

  1. npm version minor (or patch/major). The version script copies the new version into server.json and SERVER_VERSION in src/index.ts.

  2. Merge that change to main.

  3. Push a matching tag from main, for example git tag v2.1.0 && git push origin v2.1.0.

The Release workflow then:

  1. Fails unless the tag equals v + the package.json version.

  2. Runs npm ci, npm test (build, smoke tests, unit tests, OpenAPI contract), npm run verify:package (clean consumer and Node 18 transports), and npm pack --dry-run.

  3. Publishes to npm with provenance, in the npm environment. A version already on npm is skipped, so a failed run can be re-run.

  4. Waits for the version to appear on npm, then runs mcp-publisher login github-oidc and mcp-publisher publish.

One-time npm setup: on npmjs.com, open @backwork/mcp > Settings > Trusted publishing, choose GitHub Actions, and enter user tylergibbs1, repository backwork-mcp, workflow release.yml, environment npm.

Environment Variables

Variable

Required

Description

BACKWORK_API_KEY

Stdio yes; HTTP no

Backwork API key. In HTTP mode, prefer Authorization: Bearer per request.

BACKWORK_API_BASE

No

Override the API base URL.

BACKWORK_MCP_TRANSPORT

No

stdio or http.

BACKWORK_MCP_HOST

No

HTTP bind host. Defaults to 127.0.0.1.

BACKWORK_MCP_PORT

No

HTTP bind port.

BACKWORK_MCP_PATH

No

HTTP MCP path.

BACKWORK_MCP_ALLOWED_ORIGINS

No

Comma-separated allowed HTTP origins. Loopback origins are allowed for loopback requests.

BACKWORK_MCP_ALLOW_ORIGIN

No

Backward-compatible alias for BACKWORK_MCP_ALLOWED_ORIGINS.

BACKWORK_MCP_ALLOWED_HOSTS

No

Comma-separated allowed HTTP Host headers for public deployments.

BACKWORK_MCP_ALLOW_HOST

No

Backward-compatible alias for BACKWORK_MCP_ALLOWED_HOSTS.

BACKWORK_MCP_PUBLIC_HOST

No

Primary public host allowed for HTTP requests.

BACKWORK_MCP_PUBLIC_URL

No

Canonical public origin for OAuth metadata, e.g. https://backworkhealth.com.

BACKWORK_MCP_ALLOW_ENV_KEY

No

Allow private HTTP requests without bearer auth to use BACKWORK_API_KEY.

BACKWORK_MCP_AUTH_MODE

No

HTTP bearer mode: api-key, oauth, or dual. Defaults to dual when OAuth authorization servers are configured, otherwise api-key.

BACKWORK_MCP_OAUTH_AUTHORIZATION_SERVERS

OAuth

Comma-separated OAuth issuer / authorization server URLs advertised in protected-resource metadata.

BACKWORK_MCP_OAUTH_RESOURCE

No

Override the RFC 8707 resource identifier. Defaults to the public MCP URL.

BACKWORK_MCP_OAUTH_SCOPES

No

Space- or comma-separated scopes advertised to clients. Defaults to backwork:mcp.

BACKWORK_MCP_OAUTH_REQUIRED_SCOPES

No

Space- or comma-separated scopes required after token introspection.

BACKWORK_MCP_OAUTH_INTROSPECTION_URL

No

RFC 7662 token introspection endpoint used to validate OAuth access tokens.

BACKWORK_MCP_OAUTH_INTROSPECTION_CLIENT_ID

No

Client ID for introspection basic auth.

BACKWORK_MCP_OAUTH_INTROSPECTION_CLIENT_SECRET

No

Client secret for introspection basic auth.

BACKWORK_MCP_OAUTH_INTROSPECTION_TOKEN

No

Bearer token for introspection when basic auth is not used.

BACKWORK_MCP_OAUTH_API_KEY_CLAIM

No

Dot-path claim from introspection response to use as the downstream Backwork credential. If omitted, the OAuth access token is forwarded.

BACKWORK_MCP_OAUTH_EXPECTED_AUDIENCE

No

Comma-separated allowed aud values when introspection responses include an audience.

BACKWORK_MCP_EXPOSE_UNAVAILABLE_TOOLS

No

true offers tools and actions the production API marks unavailable. Write-scope actions stay hidden from read-only OAuth grants.

Troubleshooting

Missing API Key

For stdio, set BACKWORK_API_KEY in the MCP client configuration. For HTTP API-key mode, send Authorization: Bearer <key>. For HTTP OAuth mode, configure BACKWORK_MCP_OAUTH_AUTHORIZATION_SERVERS and send Authorization: Bearer <access_token>.

401 From HTTP MCP

The remote server did not receive a bearer token. Configure your MCP client to authenticate with OAuth or send an Authorization header. OAuth-enabled deployments include resource_metadata in the WWW-Authenticate header to point clients at /.well-known/oauth-protected-resource.

Claude Code OAuth

If Claude Code does not open the browser, run /mcp, select backwork, and choose the authenticate action. If it gives you a URL instead of opening a browser, copy that URL into your browser.

If the browser redirect back to Claude Code fails after consent, copy the full callback URL from the browser address bar and paste it into the Claude Code prompt.

This server does not hold OAuth tokens. It validates each request's access token by introspection and forwards it (or the mapped API key) to the Backwork API. Refreshing an expired access token is the MCP client's job: when introspection reports a token inactive, the server answers 401 with error="invalid_token", and the client can use its refresh token with the Backwork authorization server.

If Claude Code keeps using an old token, open /mcp, select backwork, clear authentication, then authenticate again. You can also remove and re-add the server with:

claude mcp remove backwork
claude mcp add --transport http --scope user backwork https://backworkhealth.com/mcp

If discovery returns 503, the Backwork web app is intentionally refusing to advertise OAuth because production signing or Redis/KV state storage is missing.

If tool calls authenticate but fail with invalid_token or invalid_target, check that BACKWORK_MCP_RESOURCE, BACKWORK_MCP_OAUTH_RESOURCE, and BACKWORK_MCP_OAUTH_EXPECTED_AUDIENCE all use:

https://backworkhealth.com/mcp

Rate Limits

Wait for the reset window or use a higher-capacity API plan.

Support

License

MIT. See LICENSE.

Claude Code plugin

The backwork plugin bundles the hosted MCP server with four research skills: prior authorization, coverage checks, policy changes, and denial appeal prep. Install it from this repository's plugin marketplace:

/plugin marketplace add tylergibbs1/backwork-mcp
/plugin install backwork@backwork

Then run /mcp, select plugin:backwork:backwork, and finish the OAuth sign-in. New accounts start with 100 free credits. See plugins/backwork/README.md for the skills and the /backwork:pa and /backwork:coverage commands.

The marketplace manifest is .claude-plugin/marketplace.json. npm test checks the manifests and every SKILL.md. CI also runs claude plugin validate --strict on the marketplace and the plugin.

Claude.ai skills

The same four skills work in claude.ai. They need the Backwork connector to call tools.

  1. Add the connector: Settings > Connectors > Add custom connector, name it Backwork, and enter https://backworkhealth.com/mcp. Finish the OAuth sign-in.

  2. Build one zip per skill. The zips go to dist/claude-skills/, which git ignores:

    node scripts/build-claude-skills.mjs

    Each zip holds the skill folder at its root, for example coverage-check.zip contains coverage-check/SKILL.md. The build fails if a SKILL.md breaks the claude.ai rules (name matches its folder, description 200 characters or fewer).

  3. Upload each zip: open Customize > Skills (in older versions, Settings > Capabilities > Skills), click Add, and select the zip. Custom skills need a Pro, Max, Team, or Enterprise plan with code execution turned on. Each user uploads their own copy.

Available Tools

8 tools
check_prior_authAInspect

Check if procedures require prior authorization for Medicare. Returns PA requirement, confidence level, matched LCD/NCD policies, and documentation checklist. Essential for determining Medicare coverage requirements before procedures.

Examples:

  • check_prior_auth(["76942"]) - check PA for ultrasound guidance

  • check_prior_auth(["76942"], { state: "TX" }) - check for Texas patient (determines MAC jurisdiction)

  • check_prior_auth(["J0585", "64493"]) - check multiple procedure codes

ParametersJSON Schema
NameRequiredDescriptionDefault
procedure_codesYesCPT/HCPCS codes to check (1-10 codes)
stateNoTwo-letter state code to determine MAC jurisdiction (e.g., TX, CA)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context by detailing the return values (PA requirement, confidence level, matched policies, documentation checklist), which helps the agent understand what to expect. However, it lacks information on potential errors, rate limits, or authentication needs, leaving some behavioral aspects unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the core purpose and key return values, followed by essential usage context and practical examples. Every sentence adds value without redundancy, making it efficient and well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of checking prior authorizations, no annotations, and no output schema, the description does a good job by explaining the purpose, usage, and return values. However, it could be more complete by detailing error handling or output structure, which would help the agent better anticipate results. It compensates well but has minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds minimal value beyond the schema by implying the purpose of parameters in examples (e.g., state determines MAC jurisdiction), but it does not provide additional syntax or format details. This meets the baseline score of 3 when schema coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Check if procedures require prior authorization for Medicare') and resource ('procedures'), distinguishing it from siblings like 'compare_policies' or 'lookup_code' by focusing on authorization requirements rather than policy comparison or code lookup. It explicitly mentions Medicare coverage, making the purpose distinct and well-defined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for usage ('Essential for determining Medicare coverage requirements before procedures') and includes examples that illustrate when to use the tool, such as checking single or multiple codes and specifying state jurisdiction. However, it does not explicitly state when not to use it or name alternatives among sibling tools, leaving some guidance implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_policiesAInspect

Compare coverage policies across different MAC jurisdictions for specific procedure codes. Useful to understand regional coverage differences for the same procedures. Shows national vs. jurisdiction-specific policies.

Examples:

  • compare_policies(["76942"]) - compare ultrasound guidance coverage nationally

  • compare_policies(["76942", "76937"], { jurisdictions: ["JM", "JH"] }) - compare specific regions

ParametersJSON Schema
NameRequiredDescriptionDefault
procedure_codesYesCPT/HCPCS codes to compare (1-10 codes)
policy_typeNoFilter by policy type
jurisdictionsNoSpecific jurisdictions to compare (e.g., ['JM', 'JH'])

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool compares policies and shows 'national vs. jurisdiction-specific policies,' adding context about output scope. However, it lacks details on behavioral traits such as rate limits, authentication needs, error handling, or what the comparison output looks like (e.g., structured data, limitations). The examples hint at usage but don't fully compensate for missing annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: the first sentence states the core purpose, followed by a usage note and output scope. The examples are concise and illustrative, adding practical value without redundancy. Every sentence earns its place, and there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description provides adequate context for a read-only comparison tool but has gaps. It explains what the tool does and includes examples, but lacks details on output format, error conditions, or limitations (e.g., max jurisdictions). For a tool with 3 parameters and no structured output, it's minimally viable but could be more complete regarding behavioral aspects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema: it mentions 'procedure codes' and 'jurisdictions' in the examples but doesn't provide additional semantics like format details or usage nuances. With high schema coverage, the baseline is 3, and the description doesn't significantly enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Compare coverage policies across different MAC jurisdictions for specific procedure codes.' It specifies the verb ('compare'), resource ('coverage policies'), scope ('across different MAC jurisdictions'), and target ('specific procedure codes'), distinguishing it from siblings like 'get_policy' or 'lookup_code' which likely retrieve individual policies or codes rather than comparative analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: 'Useful to understand regional coverage differences for the same procedures.' It implies usage for comparative analysis across jurisdictions, but does not explicitly state when not to use it or name specific alternatives among the sibling tools (e.g., 'get_policy' for single policies). The examples help illustrate usage but don't provide explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_policyAInspect

Get detailed information about a specific Medicare coverage policy. Use this after finding a policy ID from search_policies or lookup_code. Can include criteria, codes, attachments, and version history.

Examples:

  • get_policy("L33831") - LCD for ultrasound guidance

  • get_policy("A52458", { include: "criteria,codes" }) - with coverage criteria

ParametersJSON Schema
NameRequiredDescriptionDefault
policy_idYesPolicy ID (e.g., L33831, A52458, NCD220.6)
includeNoAdditional data: 'criteria', 'codes', 'attachments', 'versions'

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It describes what information is included (criteria, codes, attachments, version history) but doesn't disclose behavioral aspects like rate limits, authentication requirements, error handling, or response format. The examples help illustrate usage but don't add behavioral context beyond what's already stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, followed by usage guidance and examples. Every sentence earns its place: the first states what it does, the second when to use it, the third details included information, and the examples illustrate parameter usage. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description does well by specifying the tool's purpose, usage context, and included data types. However, it lacks details on behavioral traits (e.g., response structure, error cases) and doesn't fully compensate for the missing output schema, though the examples provide some insight into returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (policy_id and include) with their types and descriptions. The description adds minimal value by showing example usage with the parameters, but doesn't provide additional semantics beyond what the schema provides, such as format details for policy_id or how 'include' values combine.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and resource 'detailed information about a specific Medicare coverage policy', distinguishing it from siblings like search_policies (which finds policies) and lookup_code (which finds codes). It specifies the exact scope of information retrieved: criteria, codes, attachments, and version history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: 'Use this after finding a policy ID from search_policies or lookup_code.' It provides clear context for usage and distinguishes it from alternatives by specifying it's for detailed information on a specific policy, not for searching or comparing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_policy_changesBInspect

Track recent changes to Medicare coverage policies. Useful for monitoring updates, new policies, and retirements. Can filter by date, policy ID, or change type.

Examples:

  • get_policy_changes() - recent changes

  • get_policy_changes({ since: "2024-01-01T00:00:00Z" }) - changes since date

  • get_policy_changes({ policy_id: "L33831" }) - changes to specific policy

ParametersJSON Schema
NameRequiredDescriptionDefault
sinceNoISO8601 timestamp - only changes after this date
policy_idNoFilter to a specific policy
change_typeNoFilter by type of change
limitNoResults per page
cursorNoPagination cursor

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions tracking changes and filtering, but fails to disclose key behavioral traits: it doesn't specify if this is a read-only operation, what permissions are needed, whether there are rate limits, or the format of returned data. The examples hint at pagination ('limit' parameter) but don't explain it. For a tool with no annotations, this leaves significant gaps in understanding how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: the first sentence states the purpose clearly, followed by brief usage notes and helpful examples. The examples are relevant but could be more integrated; overall, there's little wasted text, and the structure supports quick understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 parameters, no output schema, no annotations), the description is moderately complete. It covers purpose and basic usage but lacks details on behavioral aspects (e.g., safety, permissions) and output format. Without annotations or output schema, the agent must infer behavior from the description alone, which is insufficient for full transparency. It's adequate but has clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema: it lists filtering options ('date, policy ID, or change type') and provides usage examples, but doesn't explain parameter interactions or add semantic context not in the schema. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Track recent changes to Medicare coverage policies' with specific verbs ('track', 'monitor') and resources ('changes', 'policies'). It distinguishes from siblings like 'get_policy' (single policy) and 'search_policies' (searching rather than tracking changes), though not explicitly. The purpose is specific but could be more explicit about sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implied usage through examples (e.g., 'Useful for monitoring updates') and parameter examples, but lacks explicit guidance on when to use this tool versus alternatives like 'get_policy' or 'search_policies'. It mentions filtering capabilities but doesn't specify scenarios where this tool is preferred over siblings, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_jurisdictionsAInspect

Get list of Medicare Administrative Contractor (MAC) jurisdictions. Returns MAC names, jurisdiction codes, and covered states. Use this to find the right jurisdiction for a patient's state.

Example:

  • list_jurisdictions() - get all MAC jurisdictions and their states

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It describes the return data (MAC names, codes, states) and includes an example, but lacks details on behavioral traits such as rate limits, error handling, or data freshness. The description adds some value but doesn't fully compensate for the absence of annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: it starts with the core purpose, then details the output, usage guidance, and an example. Every sentence adds value without redundancy, making it efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is complete enough for a read-only lookup tool. It explains what it does, what it returns, and how to use it, though it could benefit from more behavioral context like response format or limitations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, and schema description coverage is 100%, so no parameter documentation is needed. The description doesn't add parameter semantics, but with no parameters, a baseline of 4 is appropriate as it doesn't need to compensate for any gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Get list'), resource ('Medicare Administrative Contractor (MAC) jurisdictions'), and output details ('MAC names, jurisdiction codes, and covered states'). It distinguishes itself from sibling tools like 'check_prior_auth' or 'get_policy' by focusing on jurisdiction lookup rather than policy or authorization operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: 'Use this to find the right jurisdiction for a patient's state.' This provides clear context for application, though it doesn't specify when not to use it or name alternatives among siblings, but the guidance is sufficient for a 5 given the explicit purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_codeAInspect

Look up a medical code (CPT, HCPCS, ICD-10, or NDC) and get coverage information. Returns code details, descriptions, RVU values, and related Medicare policies. Use this to understand what a code means and whether it's covered.

Examples:

  • lookup_code("76942") - ultrasound guidance

  • lookup_code("J0585") - Botox injection

  • lookup_code("M54.5") - low back pain diagnosis

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe medical code to look up (e.g., 76942, J0585, M54.5)
code_systemNoCode system hint - auto-detected if not provided
jurisdictionNoMAC jurisdiction code to filter policies (e.g., JM, JH)
includeNoAdditional data: 'rvu', 'policies', or 'rvu,policies'
fuzzyNoEnable fuzzy matching for typos/partial codes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does well by stating what the tool returns (code details, descriptions, RVU values, Medicare policies) and providing concrete examples. However, it doesn't mention important behavioral aspects like rate limits, authentication requirements, error handling, or whether this is a read-only operation (though implied by 'look up').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly structured: a clear purpose statement, followed by what it returns, then usage guidance, and finally concrete examples. Every sentence adds value with zero waste. The examples are relevant and illustrative without being verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, 100% schema coverage, but no annotations or output schema, the description does quite well. It explains the tool's purpose, what it returns, and when to use it. The examples provide concrete usage patterns. The main gap is the lack of output format details (since no output schema exists), but the description compensates by listing the types of information returned.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all 5 parameters. The description adds minimal parameter semantics beyond the schema - it mentions the types of codes (CPT, HCPCS, ICD-10, NDC) which aligns with the code_system enum, and the examples show code formats. This meets the baseline of 3 when schema coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Look up a medical code... and get coverage information.' It specifies the exact action (look up), resource (medical codes), and output (coverage info, code details, descriptions, RVU values, Medicare policies). The examples reinforce this by showing specific code lookups, distinguishing it from sibling tools like check_prior_auth or search_policies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: 'Use this to understand what a code means and whether it's covered.' This gives practical guidance on its primary use case. However, it doesn't explicitly mention when NOT to use it or name specific alternatives among the sibling tools, which prevents a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_criteriaAInspect

Search through coverage criteria blocks across Medicare policies. Find specific indications, limitations, or documentation requirements. More targeted than full policy search.

Examples:

  • search_criteria("diabetes") - criteria mentioning diabetes

  • search_criteria("BMI", { section: "indications" }) - BMI requirements for coverage

  • search_criteria("frequency", { section: "limitations" }) - frequency limitations

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch query for criteria text
sectionNoFilter by criteria section type
policy_typeNoFilter by policy type
jurisdictionNoFilter by MAC jurisdiction
limitNoResults per page
cursorNoPagination cursor

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a read-only search operation without stating it explicitly, and it doesn't cover aspects like rate limits, authentication needs, or pagination behavior (beyond the cursor parameter in the schema). The examples add some context but lack comprehensive behavioral details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with a clear purpose statement followed by targeted examples. Every sentence earns its place by reinforcing usage or providing practical guidance, with no wasted words or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a search tool with 6 parameters and no output schema, the description is reasonably complete for guiding usage but lacks details on return values or error handling. It compensates well with examples and context, though it could be more comprehensive for a tool with multiple filtering options.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema by mentioning 'section' in examples, but it doesn't provide additional meaning, syntax, or format details for parameters like 'policy_type' or 'jurisdiction' that aren't covered in the examples.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('Search through coverage criteria blocks') and resources ('across Medicare policies'), distinguishing it from siblings like 'search_policies' by emphasizing it's 'more targeted than full policy search' and focuses on 'specific indications, limitations, or documentation requirements'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('more targeted than full policy search'), but it doesn't explicitly state when not to use it or name specific alternatives among the sibling tools, such as 'search_policies' for broader searches or 'get_policy' for full policy retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_policiesAInspect

Search Medicare coverage policies (LCDs, NCDs, Articles). Use this to find policies related to procedures, conditions, or coverage questions. Supports keyword and semantic search modes.

Examples:

  • search_policies("ultrasound guidance") - find policies about ultrasound

  • search_policies("diabetes CGM") - find continuous glucose monitor policies

  • search_policies("", { policy_type: "NCD" }) - list all National Coverage Determinations

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoSearch query - leave empty to browse
modeNoSearch mode: keyword (exact) or semantic (conceptual)keyword
policy_typeNoFilter by policy type
jurisdictionNoMAC jurisdiction code (e.g., JM, JH, JK)
payerNoFilter by payer name
statusNoPolicy status filteractive
limitNoResults per page
cursorNoPagination cursor from previous response
includeNoAdditional data: 'summary', 'criteria', 'codes'

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. While it mentions search modes and gives examples, it doesn't describe important behavioral traits: whether this is a read-only operation, what the response format looks like (e.g., list of policy summaries), pagination behavior beyond the cursor parameter, rate limits, authentication requirements, or error conditions. For a search tool with 9 parameters and no annotations, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and appropriately sized. It starts with the core purpose, provides usage context, mentions search modes, and gives three helpful examples. Every sentence earns its place, and the examples are directly relevant to demonstrating tool usage without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (9 parameters, no annotations, no output schema), the description is incomplete. While it covers the basic purpose and usage, it lacks crucial information about what the tool returns (no output schema means the description should explain response format), behavioral constraints, and error handling. For a search tool with many filtering options, users need to understand what results look like and how to interpret them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 9 parameters thoroughly. The description adds minimal value beyond the schema: it mentions 'keyword and semantic search modes' (covered by the mode parameter) and gives examples that show query usage and policy_type filtering. However, it doesn't provide additional semantic context for parameters like jurisdiction, payer, or include beyond what's in the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Search Medicare coverage policies (LCDs, NCDs, Articles).' It specifies the resource (Medicare coverage policies) and the action (search), and distinguishes it from siblings like 'get_policy' (retrieve specific policy) or 'compare_policies' (compare multiple policies). The examples reinforce the search functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: 'to find policies related to procedures, conditions, or coverage questions.' It mentions search modes (keyword/semantic) which helps guide usage. However, it doesn't explicitly state when NOT to use it or when to prefer alternatives like 'get_policy' for retrieving a specific known policy or 'search_criteria' for searching within policy criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv1.0.0
    • First observedcheck_prior_auth
    • First observedcompare_policies
    • First observedget_policy
    • First observedget_policy_changes
    • First observedlist_jurisdictions
    • First observedlookup_code
    • First observedsearch_criteria
    • First observedsearch_policies

TDQS

A4/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a clearly distinct purpose within the Medicare coverage domain. For example, check_prior_auth focuses on authorization requirements, compare_policies on regional differences, get_policy on detailed policy information, and search_policies on policy discovery, with no significant overlap in functionality. The descriptions clearly differentiate their roles, making it easy for an agent to select the right tool.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern using snake_case, such as check_prior_auth, compare_policies, get_policy, and search_policies. This uniformity enhances readability and predictability, allowing agents to easily understand and navigate the toolset without confusion from mixed naming conventions.

Tool Count5/5

With 8 tools, this server is well-scoped for its purpose of Medicare coverage and policy management. Each tool serves a specific, necessary function, from checking prior authorizations to searching policies and comparing jurisdictions, providing comprehensive coverage without being overly sparse or bloated. The count aligns perfectly with the domain's complexity.

Completeness5/5

The toolset offers complete coverage for Medicare policy workflows, including discovery (search_policies, lookup_code), detailed retrieval (get_policy, get_policy_changes), comparison (compare_policies), jurisdiction handling (list_jurisdictions), and specific checks (check_prior_auth, search_criteria). There are no obvious gaps; agents can perform end-to-end tasks from code lookup to authorization assessment.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides real-time access to medical data including drug interactions, ICD-10 codes, FDA adverse event reports, and clinical guidelines. It enables LLMs to query databases like openFDA, PubMed, and CMS for pharmaceutical and clinical information.
    1
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Healthcare billing AI for agents — 12 tools for ICD-10/CPT/HCPCS code lookup (80K+ codes), prior auth prediction, medical NER, claims validation, HIPAA compliance auditing, and provider/drug enrichment. Pay-per-call via credits or USDC.
    20
    49 npm
    3
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Provides AI agents with instant access to 10M+ OMOP medical vocabulary concepts for searching, mapping, and navigating clinical codes across SNOMED, ICD-10, RxNorm, LOINC, and more.
    11
    83 npm
    6
    MIT