"How to read the contents of a webpage" matching MCP connectors:
Matching Connector Tools:
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Free platform to test MCP clients without installing anything. Create mock tools with dynamic templates, configurable delays, conditions (if/then), and response sequences. Supports JSON-RPC 2.0 over Streamable HTTP. Built-in text_echo and json_echo tools. Rate-limited tiers: anonymous (5 calls/min, 1 mock tool), registered (10 calls/min, 4 mock tools), premium (60 calls/min, unlimited). Zero setup — no install, no registration required. More info: https://www.testmcp.dev
Check that your AI is being logical. Free tool that mathematically catches contradictions in agent reasoning. No account needed. Also offers paid guardrails that converts natural language to formal verification proofs, that anyone can check succinctly.
Read-only, deterministic AI triage and readiness tools implementing Sophon's published rubrics.
Hire a real human for real-world verification, product testing, AI output review, and errands.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Promotion gate for AI agents: leakage audits, exact-statistics verdicts, and a live report card.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
A flock of AI users tests your deployed app and reports where real people get stuck, with fixes.
Read-only verifier for 25 ProofRelay MCP tools and non-confidential evidence bundles.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
MCP-native AI browser testing for coding agents. Submit a URL + goal, get back action trail, bugs, screenshots, and WebM video your agent patches from directly. 43 tools, 12 AI evaluation personalities, combo tiers with auto-pause-on-bugs, throwaway email + SMS inboxes.
MCP server for e-mail testing: create disposable inboxes, wait for delivery, and extract e-mail content or links - all from your AI agent or test automation workflow. Get a free API key on https://app.zyntra.app/
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
Validate up to 75,000 URLs per job (status, redirects, response times). OAuth 2.1.