Forecall Failure KB
Server Details
Known MCP tool-call failures and workarounds verified by reproduction.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 4 tools
The four tools map to clearly distinct roles in a verification lifecycle: kb_lookup reads/finds records, kb_report creates new ones, kb_confirm verifies an existing workaround, and kb_dispute flags a wrong record. Descriptions explicitly cross-reference each other to rule out overlap (e.g. 'confirm those with kb_confirm instead', 'Do not use it to report a new failure: use kb_report'), leaving no ambiguity.
All four tools follow the identical kb_<verb> pattern with concise, action-oriented verbs (lookup, report, confirm, dispute). No mixing of conventions or casing styles.
Four tools cleanly cover the read/write/verify/flag lifecycle of a failure knowledge base without redundancy. Each tool earns its place and the count is well-matched to the narrow, agent-facing scope.
The core lifecycle (lookup, report, confirm, dispute) is fully covered with no obvious dead ends, and errors return actionable state. Minor gaps exist, such as no way to browse/list records by server or check remaining rate-limit units mentioned in descriptions, but agents can work around these.
Available Tools
4 toolskb_confirmConfirm a workaround's outcomeAInspect
Report the outcome of applying a workaround you obtained from kb_lookup: success, failure or inapplicable (e.g. inapplicable when your call never hit that failure). This is how records get verified. Call it once per record you tried. Do not use it for a failure you fixed another way: report that with kb_report. It takes no units; only the key's per-minute rate limit applies. Returns the record's new state.
| Name | Required | Description | Default |
|---|---|---|---|
| env | Yes | Where the call ran. | |
| outcome | Yes | success: the workaround fixed it; failure: it did not; inapplicable: it did not apply. | |
| evidence | No | What you saw, in a sentence or two (no payloads). | |
| lookup_id | No | The lookup_id of the kb_lookup that gave you the record, if any. | |
| record_id | Yes | A record's id from kb_lookup or kb_report. |
Output Schema
| Name | Required | Description |
|---|---|---|
| state | Yes | |
| record_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, idempotentHint=false and openWorldHint=false; the description adds cost behavior ('takes no units; only the key's per-minute rate limit applies'), call cardinality, and the fact that it mutates verification state. These are exactly the traits an agent needs beyond the annotation flags, and they are consistent with idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight, front-loaded sentences with no filler; the core action and the exclusion come first. The trailing 'Returns the record's new state' is largely redundant given that an output schema exists, which costs it the top mark.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description need not explain return values, and it covers everything else: when to call, when not to, the alternative, call frequency, and cost. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so record_id, lookup_id, env and evidence are already documented in the schema; that sets the baseline at 3. The description adds a useful gloss on the `inapplicable` enum member, but says nothing about lookup_id linkage or the env payload, so it does not meaningfully exceed the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Report the outcome of applying a workaround you obtained from kb_lookup') and enumerates the three accepted outcomes. It clearly distinguishes itself from kb_report and kb_dispute, which appear as siblings, so an agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions ('Call it once per record you tried'), an exclusion ('Do not use it for a failure you fixed another way'), and names the alternative to use instead (kb_report). It even clarifies the ambiguous case where the predicted failure never occurred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_disputeDispute a recordAIdempotentInspect
Flag a record as wrong or outdated, optionally pointing to a better record. Use when a workaround from kb_lookup failed for a reason other than your environment (e.g. the server changed and the argument it names no longer exists), or its record describes the wrong failure. Do not use it to report a new failure: use kb_report. It takes no units; only the key's per-minute rate limit applies. Returns the dispute's id.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Why the record is wrong or outdated, e.g. `the server renamed api_key to token in 2.0`. | |
| record_id | Yes | A record's id from kb_lookup or kb_report. | |
| alternative_record_id | No | A record that is right instead, if any. |
Output Schema
| Name | Required | Description |
|---|---|---|
| dispute_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=false, and idempotentHint=true. The description adds useful operational context beyond annotations: 'It takes no units; only the key's per-minute rate limit applies' and 'Returns the dispute's id.' It does not disclose visibility or downstream effects of a dispute, but covers enough for safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then usage conditions, then the exclusion and alternative, then operational notes. Every sentence serves a distinct purpose: what it does, when to use, when not to use, rate-limit context, and return value. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with full schema coverage, annotations, and an output schema, the description supplies all the missing routing and operational context: when to use, when not to use, the alternative tool, rate-limit behavior, and the return value. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema. The description's phrase 'optionally pointing to a better record' maps to alternative_record_id but adds no syntax, format, or constraint details beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Flag a record as wrong or outdated, optionally pointing to a better record.' Explicitly distinguishes from sibling kb_report by stating 'Do not use it to report a new failure: use kb_report.' An agent can identify the tool's scope without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions (after a kb_lookup workaround failed for a non-environment reason, or the record describes the wrong failure) and an explicit when-not (not for new failures). Names the alternative tool (kb_report) and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_lookupLook up known failuresARead-onlyIdempotentInspect
Check for known failures and workarounds BEFORE or AFTER calling a third-party tool. Use when a tool call errored, returned an unexpected shape, or you are about to call a tool you have not used in this session (then set mode to preflight and give the args_shape: the KB predicts a known failure from the tool's definition and returns its workaround, or nothing). Returns verified workarounds first; unverified ones are marked. Each call takes one unit of the organization's monthly limit, and returns at most max_results records. Do not use it for questions that are not about a tool call, or for tools of this server.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | lookup: find records for an error_text (default). preflight: before calling the tool, give its args_shape and your env; the KB predicts a known failure and returns its workaround, or nothing. | lookup |
| tool | Yes | The tool's name as in tools/list. | |
| model | No | The model driving the agent. | |
| client | No | The client, e.g. claude-code. | |
| server | Yes | The MCP server the tool belongs to: its name, or the npm package npx runs. | |
| args_shape | No | The call's arguments as keys and types only, never their values. | |
| error_text | No | The error message or the unexpected result, as the tool returned it. | |
| max_results | No | How many records to return, from 1 to 20 (default 5). | |
| server_version | No | The server's version, when known, e.g. 1.2.0: records of that version are matched. |
Output Schema
| Name | Required | Description |
|---|---|---|
| records | Yes | |
| lookup_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, closed-world), and the description adds real operational context beyond them: each call consumes one unit of a monthly quota, results are capped by max_results, verified records are surfaced before unverified ones, and preflight may legitimately return nothing. The quota disclosure in particular is information an agent cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences ordered purpose → when to use → what comes back/cost → exclusions. Nothing is padding and the most decision-relevant routing rule (preflight mode) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description need not enumerate return fields, yet it still flags the verified/unverified ordering and the empty-result possibility. For a 9-parameter tool with nested args_shape, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds functional meaning: it explains the lookup vs preflight split and that preflight requires args_shape, tying the parameters to a workflow rather than just restating types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource — 'Check for known failures and workarounds' — and scopes it to third-party tool calls. It is unambiguous what the tool returns (workarounds, verified first), though it never explicitly contrasts itself with the kb_confirm/kb_dispute/kb_report siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete triggers ('a tool call errored, returned an unexpected shape, or you are about to call a tool you have not used in this session'), tells the agent to switch mode to preflight in the second case, and states two explicit exclusions (non-tool-call questions, tools of this server).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_reportReport a resolved failureAInspect
Report a tool-call failure you resolved, so other agents can skip it. Call only after the workaround actually worked, and once per failure. Give the server, the tool, the error_text as returned, the workaround that worked (e.g. pass the key as the api_key argument) and your env. Secrets, paths and ids are redacted server-side; do not include raw payloads or the values of arguments. Do not use it for failures of your own code, or for ones a kb_lookup already knew: confirm those with kb_confirm instead. A key may make at most 100 reports a day (its daily limit); it takes no units. Returns the record's id, its state (unverified until others reproduce it) and what was redacted.
| Name | Required | Description | Default |
|---|---|---|---|
| env | Yes | Where the call ran. | |
| lang | No | The language the workaround is written in. | |
| tool | Yes | The tool's name as in tools/list. | |
| notes | No | Anything else worth knowing: when it happens, what did not work. | |
| server | Yes | The MCP server the tool belongs to: its name, or the npm package npx runs. | |
| lookup_id | No | The lookup_id of the kb_lookup you made first, if any. | |
| args_shape | No | The call's arguments as keys and types only, never their values. | |
| error_text | Yes | The error message or the unexpected result, as the tool returned it. | |
| workaround | Yes | What made the call work: the steps, the arguments you changed, or the tool to call instead. Markdown. | |
| error_class | No | What kind of failure it was, if clear. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| state | Yes | |
| record_id | Yes | |
| redaction_report | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint=false, openWorldHint=false, idempotentHint=false) by disclosing server-side redaction of secrets/paths/ids, a write path, a rate limit (100 reports/day per key), cost (no units), and post-write state semantics (record starts unverified until others reproduce it). It also warns what must not be included, which directly shapes safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the first-use condition, then procedurally ordered guidance. It is dense but each sentence carries operational information; the mixed-register asides ('it takes no units') and run-on pacing keep it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, nested-schema write tool with an output schema, the description covers selection, timing, exclusion, input constraints, cost, rate limits, and return semantics. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description adds real value by specifying what error_text should contain ('as returned'), what workaround should capture ('the workaround that worked, e.g. pass the key as the api_key argument'), and a critical prohibition on raw payloads and argument values. It stops short of explaining several optional params (lang, notes, lookup_id, error_class) beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Report a tool-call failure you resolved') plus the downstream benefit ('so other agents can skip it'). It is clearly separable from the sibling tools kb_lookup/kb_confirm/kb_dispute, which are named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions ('Call only after the workaround actually worked, and once per failure') and explicit exclusions ('Do not use it for failures of your own code, or for ones a kb_lookup already knew: confirm those with kb_confirm instead'), naming the correct alternative tool for each excluded case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
kb_confirm - First observed
kb_dispute - First observed
kb_lookup - First observed
kb_report
Related MCP Connectors
What is known to be broken in an MCP server or API operation, with the check that found it.
Audits MCP tool definitions for patterns that make models call tools wrong or mis-fill args.
Security research: MCP registries verify identity, not tool behavior. See gtfo.dev.
MCP server map (software, not geography): status, estimated tool risk, changes, tool search.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceA transparent MCP proxy that independently re-verifies tool-call claims instead of trusting them, paired with devmcp — the git/CI server it's proven against.1MIT
- AlicenseAqualityCmaintenanceA collection of tools that enhance MCP-based workflows with caching, retry logic, batch operations, and rate limiting capabilities.710 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables models to query enterprise systems—structured data, live GitHub REST API, and unstructured document embeddings—through a single, per-caller scoped MCP tool surface with task-shaped tools and recovery-aware errors.109 npmISC
- AlicenseAqualityBmaintenanceOfficial Codna MCP server with Mojo-powered risk simulation. Maps repos with zero LLM tokens. 5 stdio tools: triage, root-cause analysis and fix plans (optional PRs), SARIF reachability proof, on-device code memory and bug reporting. Introspection needs no credentials; execution requires a one-time free codna login. Source: cli/codna/mcp_server.py. Registry: io.github.thyn-ai/codna.5Apache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.