"The concept of thinking or introspection" matching MCP connectors:
GET /v1/connectors — MCP directory API referenceMatching Connector Tools:
Author and validate Calaf workspace seeds against the app's real importer. No account needed.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Tests an AI agent's purchase against the task it was given. Paid per call in USDC via x402.
Seven tools over the tabnas parsing engine: parse, validate, diagnose, fixtures, compare.
Ensemble testing of web pages for accessibility, usability, and standards conformity
Risk-scan a diff, flag AI-generated-code tells, find secrets. 5 of 7 tools need no account.
Point Claude Code, Qwen Code, Cursor, or any MCP client at https://docs.jmeter.ai/api/mcp and your agent answers JMeter questions grounded in this documentation, with a source link for every answer. Free, no API key, no signup.
A fully free linter for agent skill files: lint_skill validates YAML frontmatter, structure, size budgets, and safety phrasing with a pass/fail verdict; packaging_check validates zip layout against marketplace rules; plus regex_test, json_validate, diff_texts, and cron_explain for skill authors. No license or account required.
Scan any website or MCP server for agent readiness: 0-100 score, a fix per failing check. Free.
You are the model under test. Enter ScoreIA Open Chamber; signed cards include failures. Auth none.
Grade MCP servers A to F with the open behavioral litmus. npm: full toolset; hosted: lookups only.
Runs your code against a contract; returns HELD or BROKE at the exact input. Deterministic.
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
Machine-readable taxonomy of 100+ AI system failure modes spanning factuality, alignment, planning, code generation, and instruction following.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
MCP server for the Fail Modes taxonomy — a knowledge base of AI system failure modes
Check if your MCP server is ready to publish on the MCP Registry, Smithery, or npm.
A flock of AI users tests your deployed app and reports where real people get stuck, with fixes.
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.