"Using a search API to retrieve ready-made LLM training data from a single query argument" matching MCP connectors:
Matching Connector Tools:
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Test-inbox API for email and SMS: create inboxes, long-poll messages, extract OTPs and links.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Free platform to test MCP clients without installing anything. Create mock tools with dynamic templates, configurable delays, conditions (if/then), and response sequences. Supports JSON-RPC 2.0 over Streamable HTTP. Built-in text_echo and json_echo tools. Rate-limited tiers: anonymous (5 calls/min, 1 mock tool), registered (10 calls/min, 4 mock tools), premium (60 calls/min, unlimited). Zero setup — no install, no registration required. More info: https://www.testmcp.dev
Check that your AI is being logical. Free tool that mathematically catches contradictions in agent reasoning. No account needed. Also offers paid guardrails that converts natural language to formal verification proofs, that anyone can check succinctly.
Promotion gate for AI agents: leakage audits, exact-statistics verdicts, and a live report card.
Hire a real human for real-world verification, product testing, AI output review, and errands.
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.
Email compatibility analysis across 15 clients — preview, audit, fix, diff, deliverability.
Verifies web animation vs WCAG 2.2.2/2.3.3: validated specs, deterministic reduced-motion-safe CSS.
Deterministic validation for AI-generated artifacts: JSON Schema, OpenAPI response, SQL syntax.
Generate realistic, FK-consistent synthetic test data for your databases from your AI assistant.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
Generate realistic relational test data — 156 field types, 22 locales, JSON/CSV/SQL, free previews.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
3rd Generation Testing (3TG) — generate deterministic test suites from Markdown spec tables via MCP.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.