"A list of all remote servers" matching MCP connectors:
Matching Connector Tools:
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 145+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Score any URL against a real design contract — 40 checks, A-F grade, token + motion validation.
MCP tool observatory: do registry servers answer, and are their answers true? No key.
AI integrity standards, benchmarks, and EU AI Act-aligned Tier 0 model certification
Post-scrape data cleaner, no LLM: repairs mojibake, HTML, invisible chars. Plus a verdict.
Exposes FEDLIN's public security scanners as agent-callable tools over Streamable HTTP.
A webhook inbox for agents: one call returns a live URL. Mock, verify, inspect and replay.
Preflight QA for AI-agent deliverables with structured verdicts and repair guidance.
Accessibility pre-checks (WCAG/BFSG) in a real browser + statement drafts. Pay per call.
Find MCP servers and check whether they actually respond, via live handshake probes.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Hire a real human for real-world verification, product testing, AI output review, and errands.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
Scan URLs for WCAG 2.1 violations, generate AI fixes, and produce VPAT 2.5 compliance reports.
A flock of AI users tests your deployed app and reports where real people get stuck, with fixes.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.