"An exploration of Google, the popular search engine." matching MCP connectors:
GET /v1/connectors — MCP directory API referenceMatching Connector Tools:
Author and validate Calaf workspace seeds against the app's real importer. No account needed.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Tests an AI agent's purchase against the task it was given. Paid per call in USDC via x402.
Seven tools over the tabnas parsing engine: parse, validate, diagnose, fixtures, compare.
Ensemble testing of web pages for accessibility, usability, and standards conformity
Risk-scan a diff, flag AI-generated-code tells, find secrets. 5 of 7 tools need no account.
A skeptical senior-engineer code reviewer over MCP: risk-scans unified diffs, flags AI-generated-code tells, reports complexity hotspots, scans for leaked secrets, and runs an OWASP security pass — real analyzers, no external APIs. Free tier, no signup.
Verify an agent's advertised route, price, payment details, and schemas against its live endpoint.
AgentReady.market audit: can an AI shopping agent find, understand and BUY on this store? /100.
Resolve whether an uncertain side-effecting action completed before software retries it.
You are the model under test. Enter ScoreIA Open Chamber; signed cards include failures. Auth none.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Deterministic recipe verification engine — validates AI-generated recipes against master SOPs.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
みんなのテスター(スキマラボ)公開情報MCP。Google Playクローズドテストのテスター集めを支援する相互テストコミュニティの情報を提供する。
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.