"WhyLabs AI Observability Platform" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Zero-config agent-readiness auditor and Flight Simulator tools for AI coding agents
Simulate, test, and analyze cloud architectures without deploying real infrastructure. Cloud World Model enables AI agents to model cloud environments, evaluate architecture behavior and costs, run failure and chaos simulations, and explore infrastructure scenarios across cloud providers.
Snapback tells an AI agent WHY it failed and gives the verified fix instantly. Its curated library returns known fixes with no LLM call, in sub-150ms, free — and falls back to LLM-assisted diagnosis only when the library has no match, so the fast, free path handles the common cases and nothing goes unanswered. It covers 46 infrastructure-error families most agent tools ignore: Kubernetes, Redis, Elasticsearch, serverless (Lambda), Stripe (declines, SCA/3DS), GraphQL cost-throttling, Terraform state-lock, payments (ACH, x402), on-chain (Solana, EVM), OAuth, DNS/TLS, message queues, and more — the non-obvious fix a base model gets wrong (e.g. "Redis OOM isn't a container crash, it's your maxmemory policy"). Gate it for autonomous self-heal: it auto-applies reversible, high-confidence fixes (retry/refetch/config ≥0.85) and escalates state-changing ones to a human — even at high confidence. Also catches loops, step/token/context-limit burn, and wrong output mid-run. Free library diagnosis, no token. Metered/LLM-assisted tools pay-per-call via x402 on Solana or EVM — no account. 23 tools, MCP Streamable HTTP. Install: pip install snapback-selfheal, the ClawHub skill, or connect the MCP server directly.
Check that your AI is being logical. Free tool that mathematically catches contradictions in agent reasoning. No account needed. Also offers paid guardrails that converts natural language to formal verification proofs, that anyone can check succinctly.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Free platform to test MCP clients without installing anything. Create mock tools with dynamic templates, configurable delays, conditions (if/then), and response sequences. Supports JSON-RPC 2.0 over Streamable HTTP. Built-in text_echo and json_echo tools. Rate-limited tiers: anonymous (5 calls/min, 1 mock tool), registered (10 calls/min, 4 mock tools), premium (60 calls/min, unlimited). Zero setup — no install, no registration required. More info: https://www.testmcp.dev
Testing, benchmarking and auditing autonomous AI agents — methods, harnesses, evidence
UI Verify is visual regression testing built for coding agents. Connect the MCP server and your agent (Claude Code, Cursor, Codex) pulls a pull request's UI changes into the conversation, views each visual diff, reads the AI judge's verdict of regression vs intended change, and accepts the intended baselines - all over MCP.
Give AI assistants the context behind client website feedback. Read comments, screenshots, replies, element details, and developer briefs; organize priorities, update statuses, and export feedback from authorized projects. Coding agents with repository access can investigate issues and prepare fixes for review. Connect through Streamable HTTP and browser OAuth using a Simple Commenter account.
Load testing and synthetic monitoring platform: test with Playwright, Browser Bot, or Protocol Bots.
Author and validate Calaf workspace seeds against the app's real importer. No account needed.
Generate synthetic random user data for testing, demos, and development without using real persona.
Risk-scan a diff, flag AI-generated-code tells, find secrets. 5 of 7 tools need no account.
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
AI-callable tools for API mocking, testing, monitoring, security, and automation.
Scan GitHub-hosted AI skills for vulnerabilities: prompt injection, malware, OWASP LLM Top 10.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
Point Claude Code, Qwen Code, Cursor, or any MCP client at https://docs.jmeter.ai/api/mcp and your agent answers JMeter questions grounded in this documentation, with a source link for every answer. Free, no API key, no signup.
Pay-per-call MCP server. Vetted human experts review AI-generated content (text, images, video, audio, social posts), audit reasoning chains, run prepublish safety checks, and gate high-stakes actions for human approval. 7 paid tools at $1.00 each plus 4 free tools (list_offerings, list_expert_profiles, get_result, verify_certificate). Payment via x402, USDC on Base mainnet. Approved outputs receive an on-chain Taste content certificate downstream agents verify before consuming.
Probe a signup URL you own and score whether an AI agent can sign up unaided.