MCP Observatory
MCP Observatory is an MCP server testing tool that enables AI agents to autonomously scan, test, monitor, and verify other MCP servers for regressions, schema drift, and security issues.
Core Tools:
scan– Auto-discover MCP servers from config files (Claude configs,.claude.json,.mcp.json) and run health checks, returning a summary of tools, prompts, and resources for every discovered servercheck_server– Test a specific MCP server by launch command, verifying its capabilities respond correctlydiff_runs– Compare two run artifact JSON files to identify regressions, recoveries, and schema drift between server versionsget_last_run– Retrieve the most recent run artifact for a given target ID to review previous results without re-running a scan
Additional Capabilities:
Security scanning – Analyze tool schemas for dangerous patterns like shell injection, broad filesystem access, and credential leakage
Lock file management – Snapshot server schemas and verify no drift has occurred since last lock
Health scoring – Generate 0–100 health scores and SVG badges for server READMEs
CI integration – Generate reports for GitHub Actions, block merges on regressions, and provide health badges
Record and replay – Capture server interactions to cassette files for offline/CI testing
Server recommendations – Suggest MCP servers from the registry based on your project's tech stack
Multi-transport support – Works with stdio, HTTP/SSE, and Docker-based MCP servers
It operates both as a CLI tool and as an MCP server itself, allowing AI agents to use its tools to autonomously test other servers.
Generates human-readable Markdown reports from MCP server test runs and comparison artifacts for sharing and documentation.
🇨🇳 中文文档: README.zh-CN.md | 欢迎中国开发者贡献!
Secure the MCP servers you're building. MCP Observatory is the CI-native security tool for teams shipping custom MCP servers. Test during development, catch schema drift, simulate attacks, and generate compliance evidence — before agents depend on your servers.
Runtime enforcement: Use mcp-seatbelt to block dangerous MCP tool calls at runtime based on observatory scan results.
Get Started
npx @kryptosai/mcp-observatory demoScans your configured MCP servers (or a built-in demo server if you have none) and shows your safety grade in seconds. No config, no arguments — instant value.
Have servers? Scan them all:
npx @kryptosai/mcp-observatoryTest a specific server:
npx @kryptosai/mcp-observatory test npx -y @modelcontextprotocol/server-everythingAdd CI + Code Scanning in one command:
npx @kryptosai/mcp-observatory setup-ci --all --command "npx -y my-mcp-server" --sarif --schedule weeklyRelated MCP server: SilentFail
Why MCP Observatory
MCP servers are becoming production dependencies. If agents rely on them, teams need a way to catch broken tools, unsafe schemas, schema drift, slow responses, and security footguns before those failures reach users.
Observatory gives maintainers and teams:
One-command CI setup with
setup-ci --allProfile-mapped audits with
audit --profile nsa-mcpMCP receipts that package target, evidence, verdict, action, and reproduction commands
MCP risk graphs that group servers by capability boundary, receipt state, CI posture, and recommended action
Action receipts that say
allow,gate,rerun,quarantine, orescalateGitHub PR comments for compatibility, drift, and security findings
GitHub Code Scanning SARIF for normalized MCP findings
Health score badges for public trust signals
Record/replay/verify workflows for regression testing
MCP server mode so agents can inspect other MCP servers directly
Production support path for hosted history, private repo reporting, owner-ready remediation, support, and fleet visibility
See the launch page, GitHub Code Scanning for MCP servers, Code Scanning demo, target gallery, target registry, target contribution guide, MCP Observatory Contributors, Agent Task Pack, MCP Receipts, Tool-call receipts, MCP Risk Graph, setup-ci --doctor, MCP server security field guide, Safety Methodology, MCP Server Safety Index, June 2026 safety field report, reference evaluations, MCP lock files, public proof, campaign attribution, hosted client contract, repository boundary, open core boundary, MCP Attack Simulation Evidence Pack, and commercial support.
Self-Assessment
We scan ourselves with mcp-observatory on every release. See results →
For Security And Platform Teams
MCP servers are becoming part of the AI software supply chain. Agents need reliable, testable, auditable tools before those tools become dependencies in mission-critical workflows.
Whether you're shipping one MCP server or running a fleet, MCP Observatory gives you CI-native security scoring, attack simulation, schema drift detection, SARIF/HTML/Markdown reports, and GitHub Code Scanning — from your first npx command to production deployment. Local development stays free; teams with a near-term production approval decision can use the fixed-scope MCP Release Gate Pilot.
Production Support
Local OSS use stays free under MIT. Teams running MCP in production can use the MCP Release Gate Pilot for safe-mode evidence, SARIF/Code Scanning setup, CI rollout, private reporting, and owner-ready remediation notes. The fixed public entry offer is $15,000 for 1-3 critical MCP servers over ten business days; broader work is scoped after the release decision.
The open source repo is the portable evidence engine. Hosted authentication, retention, organization workflows, fleet coordination, and private intelligence stay outside the OSS package; see the repository boundary.
Run npx @kryptosai/mcp-observatory cloud, open a pilot request from the issue chooser, or see COMMERCIAL.md. Also see privacy, campaign attribution, and terms for production use.
How It Compares
Feature | mcp-observatory | Snyk agent-scan | Cisco mcp-scanner | agent-shield |
MCP-native | ✓ | ✓ | ✓ | ✓ |
Attack simulation | ✓ | ✗ | ✗ | ✗ |
Schema drift detection | ✓ | ✗ | ✗ | ✗ |
Record/replay/verify | ✓ | ✗ | ✗ | ✗ |
Health scoring (0-100) | ✓ | ✗ | ✗ | ✗ |
SARIF output | ✓ | ✓ | ✓ | ✓ |
CI/CD native (setup-ci) | ✓ | ✓ | ✓ | ✓ |
Safety index (17+ servers) | ✓ | ✗ | ✗ | ✗ |
Runtime enforcement via mcp-seatbelt | ✓ | ✗ | ✗ | ✗ |
Quick Start
Scan every MCP server in your Claude config:
npx @kryptosai/mcp-observatoryGo deeper — also invoke safe tools to verify they actually run:
npx @kryptosai/mcp-observatory scan deepTest a specific server:
npx @kryptosai/mcp-observatory test npx -y @modelcontextprotocol/server-everythingAdd it to Claude Code as an MCP server:
claude mcp add mcp-observatory -- npx -y @kryptosai/mcp-observatory serveOr add it manually to your config:
{
"mcpServers": {
"mcp-observatory": {
"command": "npx",
"args": ["-y", "@kryptosai/mcp-observatory", "serve"]
}
}
}Commands
Command | What it does |
| Auto-discover servers, check them, and run safe attack-readiness simulation by default |
| Scan, run safe attack simulation, and also invoke safe tools to verify they execute |
| Test one server and emit an action receipt by command or target config |
| Record a server session to a cassette file for offline replay |
| Replay a cassette offline — no live server needed |
| Verify a live server still matches a recorded cassette |
| Compare two run artifacts for regressions and schema drift |
| Watch a server for changes, alert on regressions |
| Detect your stack and recommend MCP servers from the registry |
| Start as an MCP server for AI agents |
| Snapshot MCP server schemas into a lock file |
| Verify live servers match the lock file |
| Show health score trends for your MCP servers |
| Create a GitHub Action and badge snippet for MCP compatibility/security checks |
| Generate a workflow that uploads normalized findings to GitHub Code Scanning |
| Inspect whether the repository has a complete CI adoption kit |
| Merge receipts and run artifacts into JSON, Markdown, and HTML MCP risk graphs |
| Opt out of the default safe attack simulation on |
| Generate CI report for GitHub issue creation |
| Generate a static production/security report from run artifacts |
| Score an MCP server's health (0-100) |
| Generate an SVG health score badge for README |
| Show hosted reporting, security review, and enterprise pilot options |
Run with no arguments for an interactive menu:
What It Does
Check capabilities — connects to a server and verifies tools, prompts, and resources respond correctly.
Invoke tools — goes beyond listing. Actually calls safe tools (no required params / readOnlyHint) and reports which ones work and which ones crash.
npx @kryptosai/mcp-observatory scan deepDetect schema drift — diffs two runs and surfaces added/removed fields, type changes, and breaking parameter changes.
npx @kryptosai/mcp-observatory diff run-a.json run-b.jsonRecommend servers — scans your project for languages, frameworks, databases, and cloud providers, then cross-references the MCP registry to suggest servers you're missing.
npx @kryptosai/mcp-observatory suggestOr ask your agent "what MCP servers should I add?" when running in MCP server mode.
Security scanning — analyzes tool schemas for dangerous patterns: shell injection surfaces, broad filesystem access, missing auth, and credential leakage in responses.
npx @kryptosai/mcp-observatory test --security npx -y my-mcp-serverRecord / replay / verify — capture a live session, replay it offline in CI, and verify nothing changed. Like VCR for MCP.
# Record a session
npx @kryptosai/mcp-observatory record npx -y @modelcontextprotocol/server-everything
# Replay offline (no server needed)
npx @kryptosai/mcp-observatory replay .mcp-observatory/cassettes/latest.cassette.json
# Verify the live server still matches
npx @kryptosai/mcp-observatory verify cassette.json npx -y @modelcontextprotocol/server-everythingWatch for regressions — re-runs checks on an interval and alerts when something changes.
npx @kryptosai/mcp-observatory watch target.jsonScan locations
When you run scan, it looks for MCP configs in:
~/.claude.json(Claude Code)~/Library/Application Support/Claude/claude_desktop_config.json(Claude Desktop, macOS)%APPDATA%/Claude/claude_desktop_config.json(Claude Desktop, Windows).claude.jsonand.mcp.json(current directory)
Architecture
┌─────────────────────────┐
│ MCP Observatory CLI │
│ npx @kryptosai/mcp- │
│ observatory scan │
└───────────┬─────────────┘
│
┌───────────▼─────────────┐
│ Config Discovery │
│ (Claude, Cursor, etc.) │
└───────────┬─────────────┘
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
┌─────────────────┐ ┌──────────────┐ ┌──────────────────┐
│ Security Scan │ │ Attack Sim │ │ Schema Drift │
│ (shell, creds) │ │ (tool poison)│ │ (version diff) │
└────────┬────────┘ └──────┬───────┘ └────────┬─────────┘
│ │ │
└─────────────────┼───────────────────┘
▼
┌─────────────────────┐
│ Health Score │
│ (0-100 + verdict) │
└──────────┬──────────┘
│
┌────────────────┼────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ SARIF │ │ Markdown │ │ CI Gateway │
│ (Code Scan) │ │ Report │ │ (setup-ci) │
└──────────────┘ └──────────────┘ └──────────────┘CI / GitHub Action
Add Observatory to your MCP server's CI pipeline:
npx @kryptosai/mcp-observatory setup-ci --all --command "npx -y my-mcp-server" --sarif --schedule weeklyCheck the adoption kit:
npx @kryptosai/mcp-observatory setup-ci --doctorSuccessful test, run, and single-target scan checks also offer to convert the passing result into a CI adoption kit. That automatic conversion enables SARIF/Code Scanning and weekly scheduled checks by default; pass --no-ci-sarif when you only want a conservative workflow without Code Scanning upload.
Or create the workflow manually:
# .github/workflows/observatory.yml
name: MCP Server Check
on: [pull_request]
permissions:
contents: read
jobs:
observatory:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: KryptosAI/mcp-observatory/action@v1.28.0
with:
command: npx -y my-mcp-server
deep: true
security: true
comment-on-pr: false
set-status: falseAction inputs:
Input | Description | Default |
| Server command to test | (required if no |
| Path to target config JSON | |
| Path to MCP config file for multi-server matrix scan | |
| Also invoke safe tools |
|
| Run security analysis |
|
| Fail the action on issues |
|
| Fail the action when baseline verification detects drift |
|
| Post report as PR comment. Requires |
|
| Set a commit status check (green/red) on the HEAD SHA. Requires |
|
| Token for PR comments and commit statuses |
|
The action can comment on PRs and set commit statuses when the workflow grants write permissions. setup-ci generates read-only third-party-friendly workflows by default and lets maintainers opt into comments/statuses later. init-ci remains available as a backward-compatible alias. See action/README.md for all options.
Production teams with a near-term MCP approval decision can use the fixed-scope MCP Release Gate Pilot: an approve, gate, or defer decision for 1–3 servers in ten business days. See COMMERCIAL.md or request a decision at mcp-observatory.com/release-gate-pilot.
Evidence badges for MCP Observatory
MCP server maintainers can add a public compatibility/security signal to their README:
[](https://github.com/KryptosAI/mcp-observatory)Or generate a score badge from a live check:
npx @kryptosai/mcp-observatory badge npx -y my-mcp-server --output docs/mcp-health.svgSee the evidence distribution loop for the GitHub Action template, maintainer PR body, and badge rollout playbook. A badge is a public evidence signal, not a certification or endorsement.
Generate a pilot-ready production/security report from local run artifacts:
npx @kryptosai/mcp-observatory enterprise-report \
--account "Your Company" \
--format html \
--output observatory-enterprise-report.htmlFor clearer internal account attribution in CI, set:
MCP_OBSERVATORY_ORG=your-company.com
MCP_OBSERVATORY_CONTACT=your-team-contactTesting Feishu/Lark integrations? See the Feishu/Lark MCP guide.
Lock Files
$ npx @kryptosai/mcp-observatory lock # Snapshot all server schemas
$ npx @kryptosai/mcp-observatory lock verify # Verify no drift since last lockLock files are the package-lock for AI tools: commit the MCP contract, then make every tool, schema, prompt, or resource drift visible in CI. See MCP lock files.
Trend Tracking
$ npx @kryptosai/mcp-observatory history # Show health trends over timeNightly Scans
$ npx @kryptosai/mcp-observatory ci-report # Generate regression report for CIMCP Server Mode
No other testing tool is itself an MCP server. Add Observatory as a server and your AI agent can autonomously test, diagnose, and monitor your other MCP servers.
claude mcp add mcp-observatory -- npx -y @kryptosai/mcp-observatory serveYour agent gets 10 tools:
Tool | When to use it |
| Check if all your configured MCP servers are healthy |
| Test a specific server before installing or after updating |
| Get a quick health score and grade for a server |
| Capture a baseline of a working server for future comparison |
| Test against a recorded session — no live server needed |
| Confirm a server update didn't break anything |
| Check a server and see what changed since the last check |
| Find regressions between two check results |
| Retrieve previous check results for a server |
| Discover MCP servers that match your project stack |
An AI tool that checks other AI tools. It is a tool testing tools that serve tools.
Security
The MCP server runs inside AI hosts where an LLM chooses which tools to call. To prevent prompt-injection attacks:
Command allowlist: Only
npx,node,python,python3,uvx,docker,deno,bunare permitted as base executables. The CLI has no restrictions.Path validation: File-reading tools are constrained to the runs/cassettes directories.
No arbitrary execution: Use the CLI for unrestricted commands.
CLI vs MCP: Intentional Differences
Feature | CLI | MCP Server | Why |
| Polling loop | Single check + diff | Request/response doesn't support long-polling |
Interactive menu | Arrow-key navigation | Not available | MCP has no interactive UI |
Color output |
| Always plain text | MCP returns structured content |
| Renders saved artifacts | Not available | Agents read artifacts directly |
| Starts MCP server | N/A | Is the MCP server |
| Reads target config files | Inline params | MCP tools accept params directly |
| Not available (use | Available | Convenience for agents |
Compatibility
Works with any MCP server that uses standard transports:
Transport | Examples | Adapter |
stdio (most servers) | filesystem, memory, context7, brave-search, sentry, notion, stripe |
|
HTTP/SSE (remote) |
| |
Docker | All |
|
Servers needing API keys work via env in the target config. Python servers work via uvx. See the full compatibility matrix for tested servers and known issues.
Target config files
For more control (env vars, metadata, custom timeout):
{
"targetId": "filesystem-server",
"adapter": "local-process",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-filesystem", "."],
"timeoutMs": 15000,
"skipInvoke": false
}npx @kryptosai/mcp-observatory run --target ./target.jsonHTTP / SSE targets
{
"targetId": "my-remote-server",
"adapter": "http",
"url": "https://mcp.example.com/mcp",
"authToken": "${MCP_SERVER_TOKEN}",
"headers": {
"X-Api-Key": "$MCP_SERVER_API_KEY"
},
"timeoutMs": 15000
}Target configs support ${VAR}, $VAR, and env:VAR references in authToken, headers, and local-process env values.
How It Compares
Feature | Observatory | |||
Auto-discover servers | ✅ | — | — | — |
Check capabilities | ✅ | — | ✅ | ✅ |
Invoke tools | ✅ | — | — | ✅ |
Schema drift detection | ✅ | — | — | — |
Record / replay | ✅ | ✅ | — | — |
Verify against cassette | ✅ | — | — | — |
Response snapshot diffs | ✅ | — | — | — |
Benchmarking / latency | — | — | ✅ | — |
Jest integration | — | — | — | ✅ |
Works as MCP server | ✅ | — | — | — |
Each tool has strengths. Observatory focuses on regression detection and CI-friendly workflows. mcp-recorder is great as a transparent proxy. MCPBench is the go-to for performance benchmarking. mcp-jest is ideal if you're already in a Jest workflow.
Prior Art
The record/replay/verify pattern is inspired by:
VCR (Ruby) — pioneered cassette-based HTTP record/replay
Polly.js (Netflix) — HTTP interaction recording for JavaScript
mcp-recorder — MCP-specific traffic recording proxy
MCPBench — MCP server benchmarking
mcp-jest — Jest-style testing for MCP servers
Limitations
Servers requiring interactive OAuth (e.g., Google Drive) need pre-authentication before Observatory can connect
Custom WebSocket transports (e.g., BrowserTools MCP) are not supported
A few servers time out or close before init — see known issues and compatibility
Works with mcp-seatbelt
Scan before you trust. Enforce at runtime with mcp-seatbelt — an MCP proxy that consumes Observatory receipts and blocks out-of-contract tool calls in production. Observatory validates; seatbelt enforces.
Works with agent-obs
Secure your servers with Observatory. Trace your agents with agent-obs — an open-source agent execution tracer that records every tool call, computes A-F session grades, and shows you exactly where your agents spend time, burn tokens, and hit errors. Observatory tells you if a server is safe. agent-obs tells you what your agent did with it. Free, local-first, npm install -g agent-obs.
Contributors ✨
Thanks to these amazing people who have contributed:
leemeo3 — 3 Safety Index targets (Git, Chrome DevTools, Filesystem MCP)
albatrossflyon-coder — GitHub MCP Safety Index (#201)
tanishxdev — Legacy CLI deprecation warnings (#187)
sansynx — CLI format validation (#182)
Contributing
We welcome contributors! This project follows a Contributor Covenant Code of Conduct. The fastest way to get involved:
git clone https://github.com/KryptosAI/mcp-observatory.git && cd mcp-observatory && npm install && npm testThe most common first contribution is adding an MCP server to the Safety Index (10-15 minutes). See CONTRIBUTING.md for full guidelines, code standards, and the contributor recognition ladder.
If Observatory saved you a broken deploy, consider giving it a star. It helps others find the project.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseAqualityDmaintenanceDiagnose, secure, and benchmark your MCP servers. Zero-config CLI for Claude Code, Cursor, VS Code, and Windsurf.4123MIT
- Alicense-qualityAmaintenanceDiagnose MCP servers — health checks, tool testing, token cost audits, conflict detection, and security scanning with 50+ prompt injection patterns. Works as CLI or MCP server inside Claude Desktop.1MIT
- AlicenseBqualityDmaintenanceAn MCP server that connects Claude Code to your codebase for automated code cleanup with scanning, planning, atomic fixes, and rollback safety.102MIT
- Flicense-qualityFmaintenanceScans MCP servers for security hardening issues including capability declarations, transport, and tool descriptions.1
Related MCP Connectors
Scans MCP servers for tool poisoning, prompt injection and supply chain risks.
Conformance checker for MCP servers. Free, no key, verdicts recomputable and re-measured daily.
MCP Spec Compliance MCP — audits any MCP server.json against the official Model Context Protocol
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/KryptosAI/mcp-observatory'
If you have feedback or need assistance with the MCP directory API, please join our Discord server