"Understanding how memory is stored or memory storage methods" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Parse WebVTT, SRT, or TTML for conformance, timing, overlaps, line length, and reading speed.
Check that your AI is being logical. Free tool that mathematically catches contradictions in agent reasoning. No account needed. Also offers paid guardrails that converts natural language to formal verification proofs, that anyone can check succinctly.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
Testing, benchmarking and auditing autonomous AI agents — methods, harnesses, evidence
UI Verify is visual regression testing built for coding agents. Connect the MCP server and your agent (Claude Code, Cursor, Codex) pulls a pull request's UI changes into the conversation, views each visual diff, reads the AI judge's verdict of regression vs intended change, and accepts the intended baselines - all over MCP.
Load testing and synthetic monitoring platform: test with Playwright, Browser Bot, or Protocol Bots.
Runs your code against a contract; HELD or BROKE at the exact input. Deterministic. 0.10 USDC/call.
What is known to be broken in an MCP server or API operation, with the check that found it.
Explore Upforge services and evidence, or submit a prototype for a human-approved static fit check.
Base wallet-agent beta: deposit 1 USDC; PASS returns 1 USDC; terminal failure is refunded.
Point Claude Code, Qwen Code, Cursor, or any MCP client at https://docs.jmeter.ai/api/mcp and your agent answers JMeter questions grounded in this documentation, with a source link for every answer. Free, no API key, no signup.
Compare two versions of a JSON row list: what was added, removed or changed, field by field.
Test FIX from Claude, ChatGPT, or Cursor: free sandbox FIX sessions or your own FIXSIM instances.
Scores any public website on how usable it is by AI agents, with per-check evidence.
Check what a paid x402 endpoint or MCP server delivered, from probes anyone can repeat.
Deterministic operations reconciliation for AI agents: COMPLETE, INCOMPLETE, or NEEDS_REVIEW.
Scan any website or MCP server for agent readiness: 0-100 score, a fix per failing check. Free.
Pixel-perfect webpage screenshots rendered in a real browser, full-page or viewport, via one POST.
Check whether a URL can be opened. No browser is launched.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.