Skip to main content
Glama

mcp-agent-reliability

Pure-computation MCP server that scores AI agent trajectories, detects silent failures, loops, and reliability metrics — zero external API cost.

Available on MCPize

What problem does this solve?

AI agents often fail silently: they loop on the same tool, return empty results, or give high-confidence answers without evidence. Debugging is expensive. This MCP gives you instant scores and failure reports from any agent run log you provide.

Think of it like a doctor’s check-up for your AI agent — it looks at the “X-ray” (the run log) and tells you what is healthy and what is broken, without needing another expensive doctor (LLM).

Related MCP server: HumanProof

Tools

Tool

Description

score_trajectory_tool

0-100 reliability score + breakdown

detect_failure_modes_tool

List of loops, empty results, high-confidence-without-evidence, etc.

analyze_tool_usage_tool

Per-tool call counts, error rates

compute_success_rate_tool

Success rate across many runs

compare_trajectories_tool

Which of two runs is more reliable

Quick Start (local)

# Install
pip install -e .

# Run (HTTP on port 8080)
python -m mcp_agent_reliability.server

Or with MCP inspector / Claude Desktop / Cursor by pointing to the HTTP endpoint.

Example trajectory input

[
  {"tool": "get_weather", "status": "ok", "result": {"temp": 28}},
  {"role": "assistant", "content": "28C today", "is_final": true}
]

Why this is valuable for entrepreneurs

  • Zero running cost (no paid APIs)

  • Helps you ship reliable agents faster → happier users → more revenue

  • Can be called by the agent itself mid-run or by your CI after tests

  • Fits the “Type A” high-margin MCP pattern preferred on MCPize

Development

pytest

License

MIT

Built daily for Prince Ruhul / Prevalid by the Daily AI Project Builder.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    C
    maintenance
    Validates agent outputs in multi-agent systems to prevent coordination failures, with tools for schema verification, hallucination detection, and freshness checks, all with zero LLM cost.
    5
    24 npm
    MIT