mcp-agent-reliability
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| score_tool_descriptionA | Score how clear, specific and LLM-friendly a tool description is (0-100). Use this before adding a new tool to an agent to reduce wrong tool calls. |
| estimate_token_costA | Roughly estimate how many tokens a list of tool definitions will consume in the agent context window. Helps decide whether to enable progressive loading. |
| simulate_tool_choiceA | Given a user prompt and a list of available tools, predict which tool an agent is most likely to pick. Useful for testing tool selection before production. |
| generate_agent_testsA | Generate 3 simple test prompts that you can feed to an agent to verify it correctly selects and uses a given tool. |
| reliability_reportA | Create a short reliability report for a set of tools. Combines description scores and token estimates into one actionable summary. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool targets a distinct aspect of agent reliability: scoring descriptions, estimating token costs, simulating tool choice, generating test prompts, and producing a combined report. There is no overlap or ambiguity between them.
Most names follow a clear verb_noun pattern (score_tool_description, estimate_token_cost, simulate_tool_choice, generate_agent_tests), but 'reliability_report' deviates as a noun_noun construction. The pattern is mostly consistent with one minor deviation.
At 5 tools, the server is well-scoped for its purpose of assessing and improving MCP tool reliability. Each tool serves a distinct function without redundancy or bloat.
The set covers the core lifecycle: evaluating descriptions, estimating cost, predicting selection, generating tests, and summarizing results. A minor gap is the lack of direct test execution or runtime monitoring, but the provided surface is reasonably complete for its intended scope.