"The most widely used models" matching MCP connectors:
GET /v1/connectors — MCP directory API referenceMatching Connector Tools:
UI Verify is visual regression testing built for coding agents. Connect the MCP server and your agent (Claude Code, Cursor, Codex) pulls a pull request's UI changes into the conversation, views each visual diff, reads the AI judge's verdict of regression vs intended change, and accepts the intended baselines - all over MCP.
Author and validate Calaf workspace seeds against the app's real importer. No account needed.
Generate synthetic random user data for testing, demos, and development without using real persona.
Simulate, test, and analyze cloud architectures without deploying real infrastructure. Cloud World Model enables AI agents to model cloud environments, evaluate architecture behavior and costs, run failure and chaos simulations, and explore infrastructure scenarios across cloud providers.
Checks the structural integrity of translated resource dictionaries against a source dictionary y...
Defectbird: the site's own MCP server — dataset; every answer cites the site.
Class Q Checker: the site's own MCP server — checker, enquiry (enquiry = a human handoff, not a...
Smoke Control Checker: the site's own MCP server — checker, enquiry (enquiry = a human handoff,...
Tests an AI agent's purchase against the task it was given. Paid per call in USDC via x402.
Seven tools over the tabnas parsing engine: parse, validate, diagnose, fixtures, compare.
Post-scrape data cleaner, no LLM: repairs mojibake, HTML, invisible chars. Plus a verdict.
You are the model under test. Enter ScoreIA Open Chamber; signed cards include failures. Auth none.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
Check if your MCP server is ready to publish on the MCP Registry, Smithery, or npm.