"A code editor for developers" matching MCP connectors:
Matching Connector Tools:
Workflow planning, recovery checkpoints, coordination, fixtures, and compatibility tools for agents.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 145+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Score any URL against a real design contract — 40 checks, A-F grade, token + motion validation.
Agentic code review, no signup to try: reality gates + frontier-model review, with veto.
Can an AI shopping agent find, understand and BUY on a store? Deterministic e-commerce audit /100.
A webhook inbox for agents: one call returns a live URL. Mock, verify, inspect and replay.
Estimated game fps for any GPU or Apple Silicon chip, with the limiter and tweaks.
Pre-execution governance for AI agents. Deterministic PASS/FAIL/REVIEW verdicts, replayable proof.
Deterministic validation for AI-generated artifacts: JSON Schema, OpenAPI response, SQL syntax.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Hire a real human for real-world verification, product testing, AI output review, and errands.
Promotion gate for AI agents: leakage audits, exact-statistics verdicts, and a live report card.
Generate deterministic placeholder image URLs and packs for docs, staging, testing, and AI agents.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Generate realistic, FK-consistent synthetic test data for your databases from your AI assistant.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes