"Understanding the Concept of LLM Context" matching MCP connectors:
Matching Connector Tools:
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Verify scraped data against the live source page. Signed verdicts, $0.01 via x402.
Post-scrape data cleaner, no LLM: repairs mojibake, HTML, invisible chars. Plus a verdict.
Estimated game fps for any GPU or Apple Silicon chip, with the limiter and tweaks.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Check if your MCP server is ready to publish on the MCP Registry, Smithery, or npm.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
A flock of AI users tests your deployed app and reports where real people get stuck, with fixes.
MCP-native AI evaluation: rubric audits, eval suites, and proof reports for AI/LLM output.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
MCP server for static security analysis of Android source code
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.