"A localized knowledge base" matching MCP connectors:
Matching Connector Tools:
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Rule engine with built-in simulation. 55 MCP tools for complete business rule lifecycle management.
Hire a real human for real-world verification, product testing, AI output review, and errands.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Promotion gate for AI agents: leakage audits, exact-statistics verdicts, and a live report card.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
Scan URLs for WCAG 2.1 violations, generate AI fixes, and produce VPAT 2.5 compliance reports.
A flock of AI users tests your deployed app and reports where real people get stuck, with fixes.
Pay-per-call MCP server. Vetted human experts review AI-generated content (text, images, video, audio, social posts), audit reasoning chains, run prepublish safety checks, and gate high-stakes actions for human approval. 7 paid tools at $1.00 each plus 4 free tools (list_offerings, list_expert_profiles, get_result, verify_certificate). Payment via x402, USDC on Base mainnet. Approved outputs receive an on-chain Taste content certificate downstream agents verify before consuming.
MCP-native AI browser testing for coding agents. Submit a URL + goal, get back action trail, bugs, screenshots, and WebM video your agent patches from directly. 43 tools, 12 AI evaluation personalities, combo tiers with auto-pause-on-bugs, throwaway email + SMS inboxes.
130+ QA & dev tools for AI agents: prompt injection, RAG testing, VLM eval, guardrails. Free.
MCP server for e-mail testing: create disposable inboxes, wait for delivery, and extract e-mail content or links - all from your AI agent or test automation workflow. Get a free API key on https://app.zyntra.app/
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
Benchmark-first release surface with a read-only MCP endpoint and operator CLI.
Evaluates UI designs for WCAG accessibility issues automated scanners miss. Paid via x402 on Base.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 145+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
295k+ bug-fix patterns with MCP Hub proxy, PII filtering, and code search