"Find the LinkedIn profile of Satya Nadella" matching MCP connectors:
Matching Connector Tools:
Author and validate Calaf workspace seeds against the app's real importer. No account needed.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 290+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Verify scraped data against the live source page. Signed verdicts, $0.01 via x402.
Seven tools over the tabnas parsing engine: parse, validate, diagnose, fixtures, compare.
Risk-scan a diff, flag AI-generated-code tells, find secrets. 5 of 7 tools need no account.
Voice-powered bug reporting with 13 MCP tools. Record bugs by talking; let AI find and fix them.
AgentReady.market audit: can an AI shopping agent find, understand and BUY on this store? /100.
Find MCP servers and check whether they actually respond, via live handshake probes.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Check if your MCP server is ready to publish on the MCP Registry, Smithery, or npm.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
A flock of AI users tests your deployed app and reports where real people get stuck, with fixes.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
MCP server for static security analysis of Android source code