openclaw-output-vetter-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| verify_response_groundingA | Check that every claim in |
| find_swallowed_exceptionsA | Scan Python source code for try/except patterns that swallow errors or substitute fabricated mock data — the silent-fake-success pattern from the r/ClaudeAI thread. Flags pass-only handlers, mock-substitution returns, silent log-and-return, and bare excepts. Each finding includes a line number + severity + code excerpt. |
| review_transcriptA | Multi-turn agent transcript review — flags unverified completion claims (assistant says 'I've configured X' with no supporting tool calls), cross-turn factual contradictions, and tool calls without observable side effects. Pass an array of {role, text, tool_calls?} objects. |
| verify_action_outcomeA | v1.1+ — Compare an agent's stated outcome against actual before/after state snapshots. Catches the [@chiefofautism, 158↑] failure mode: agent runs |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| verify-this-answer | Run a grounding check on the most recent assistant answer + flag hallucinations |
| audit-this-code | Run swallowed-exception detection on a code block and explain each finding's risk + how to fix |
| verify-this-action | v1.1+ — Compare an agent's stated outcome against actual before/after snapshots; surface STATE_UNCHANGED, NO_COMMIT, TESTS_NOT_PASSING, and STATE_VIOLATED_CONSTRAINT mismatches. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| Demo: grounded answer | Sample input demonstrating a CLEAN grounding verdict |
| Demo: fabricated answer | Sample input demonstrating a FABRICATED grounding verdict |
| Demo: swallowed-exception patterns | Sample Python code with mock-substitution + pass-only + log-and-return patterns |
| Demo: action-outcome divergence (chiefofautism case) | Sample input where claim says 'I cleaned up the project structure' but before/after snapshots are identical — demonstrates the FABRICATED verdict for ACTION_OUTCOME.STATE_UNCHANGED |
TDQS
Scored across 4 tools
Each tool targets a distinct aspect of output vetting: code exception swallowing, transcript review, action outcome verification, and response grounding. No two tools overlap in purpose; descriptions clearly differentiate them.
All tool names follow a consistent verb_noun pattern in snake_case: find_swallowed_exceptions, review_transcript, verify_action_outcome, verify_response_grounding. The verbs are descriptive and the pattern is uniform.
With four tools, the server is well-scoped for its purpose. Each tool covers a critical vetting check without unnecessary bloat or gaps. The count is appropriate for a specialized vetting server.
The tool set covers the main failure modes mentioned (exception swallowing, unverified claims, action misreports, hallucinated responses). Minor gaps might include verifying tool call correctness or security issues, but the set is reasonably complete for its intended domain.