"Créer des systèmes de prompt pour les modèles de langage (LLM)" matching MCP connectors:
Matching Connector Tools:
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Evidence-gated task verification for AI agents. Decompose goals into acceptance criteria, attach proof (screenshot, curl, file), independent LLM judge accepts or rejects. 24 tools. Hosted remote MCP (streamable-http, OAuth 2.1 + DCR).
MCP-native AI evaluation: rubric audits, eval suites, and proof reports for AI/LLM output.
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
130+ QA & dev tools for AI agents: prompt injection, RAG testing, VLM eval, guardrails. Free.
The world's first named AI prompt quality score. Score, optimize, and compare LLM prompts before they hit any model. Free tier available. Built on PEEM, RAGAS, G-Eval, and MT-Bench frameworks. x402-native on Base.
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Scan GitHub-hosted AI skills for vulnerabilities: prompt injection, malware, OWASP LLM Top 10.
Remote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.
- MCP FortressOAuth
Security scanner for MCP servers. Detect vulnerabilities, prompt injection, and tool poisoning.