"The Models Resource" matching MCP connectors:
Matching Connector Tools:
Verify scraped data against the live source page. Signed verdicts, $0.01 via x402.
Estimated game fps for any GPU or Apple Silicon chip, with the limiter and tweaks.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Check if your MCP server is ready to publish on the MCP Registry, Smithery, or npm.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Check AI work against requirements and return structured verdicts, findings, and repair steps.
Run, debug, and triage tests via natural language across HyperExecute, Automation, SmartUI, and Accessibility on the TestMu AI cloud.
Evaluate, benchmark, and simulate AI agents on the VerifyAX agent-evaluation platform.
Audits MCP tool definitions for patterns that make models call tools wrong or mis-fill args.
WCAG-Compliance/wcagc-mcp lets an assistant run real axe-core accessibility scans through the user's own wcagc account, rather than guessing accessibility from markup it can see. Tools cover a single URL or PDF, a full-site crawl, saved multi-step journeys, and violation trends.
MCP server for the VerifyAX platform. Enables agent evaluation, simulation testing, and functional/non-functional verification workflows through natural language.
## Skill Catalog The library contains 42 public skills organized by Rails development concern. | Category | Examples | |----------|----------| | Planning | `create-prd`, `generate-tasks`, `plan-tickets` | | Testing | `plan-tests`, `write-tests`, `test-service`, `triage-bug` | | Code quality | `code-review`, `respond-to-review`, `security-check`, `refactor-code` | | Architecture and DDD | `define-domain-language`, `review-domain-boundaries`, `model-domain`, `review-architecture` | | Rails imple
Pluralistic human evaluation infrastructure for AI in production. Query real human reviewer verdicts on commercial AI models - scores, flag breakdowns, and side-by-side comparisons, by our community reviews.