"Setting up services on a Linux virtual machine" matching MCP connectors:
Matching Connector Tools:
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
A webhook inbox for agents: one call returns a live URL. Mock, verify, inspect and replay.
Check if your MCP server is ready to publish on the MCP Registry, Smithery, or npm.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Hire a real human for real-world verification, product testing, AI output review, and errands.
Promotion gate for AI agents: leakage audits, exact-statistics verdicts, and a live report card.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
A flock of AI users tests your deployed app and reports where real people get stuck, with fixes.
Pay-per-call MCP server. Vetted human experts review AI-generated content (text, images, video, audio, social posts), audit reasoning chains, run prepublish safety checks, and gate high-stakes actions for human approval. 7 paid tools at $1.00 each plus 4 free tools (list_offerings, list_expert_profiles, get_result, verify_certificate). Payment via x402, USDC on Base mainnet. Approved outputs receive an on-chain Taste content certificate downstream agents verify before consuming.
Benchmark-first release surface with a read-only MCP endpoint and operator CLI.
MCP-native AI browser testing for coding agents. Submit a URL + goal, get back action trail, bugs, screenshots, and WebM video your agent patches from directly. 43 tools, 12 AI evaluation personalities, combo tiers with auto-pause-on-bugs, throwaway email + SMS inboxes.
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
Validate up to 75,000 URLs per job (status, redirects, response times). OAuth 2.1.
Evaluates UI designs for WCAG accessibility issues automated scanners miss. Paid via x402 on Base.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 145+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
Evaluate, benchmark, and simulate AI agents on the VerifyAX agent-evaluation platform.
WCAG-Compliance/wcagc-mcp lets an assistant run real axe-core accessibility scans through the user's own wcagc account, rather than guessing accessibility from markup it can see. Tools cover a single URL or PDF, a full-site crawl, saved multi-step journeys, and violation trends.
AI dev tools + image generation, paid per-use with USDC on Base (x402).