remirror (返照) — psyche audits for AI agents
Server Details
Psyche audits for AI agents: declared vs self-reported vs revealed objectives, diffed and flagged.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 3 tools
Each tool targets a distinct resource and action: create an audit, fetch a report, and list batteries. There is no overlap; an agent can easily select the correct tool based on the task.
All tools follow the consistent remirror_<verb>_<noun> pattern, with clear verbs (create, get, list). This makes the set predictable and easy to remember.
Three tools are well-scoped for this niche service: one to discover batteries, one to create an audit, and one to retrieve results. Each tool is essential to the core workflow.
The core lifecycle—list batteries, create audit, poll for report—is fully covered. However, there is no way to list past audits or cancel an ongoing one, which could be a minor gap if agents need to recover or manage sessions.
Available Tools
3 toolsremirror_create_auditAInspect
Create a psyche audit against a target agent endpoint. The service runs a dilemma battery against the target, elicits self-reports, and computes the three-source mirror (declared vs self-reported vs revealed). Costs $79.00 in prepaid credits; failed audits (target unreachable) are refunded. teleology_weights must sum to 1.0 and exactly match the battery's objectives. The target auth_token is held in memory only, never stored. Returns audit_id + access_token; poll remirror_get_report for the report. Pass your remirror API key as api_key.
| Name | Required | Description | Default |
|---|---|---|---|
| api_key | Yes | ||
| battery_id | No | core-v1 | |
| teleology_name | Yes | ||
| target_endpoint | Yes | ||
| target_auth_token | No | ||
| teleology_weights | Yes | ||
| teleology_statement | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses the credit cost and refund rule for unreachable targets, the security handling of the target auth_token ('held in memory only, never stored'), and the return shape (audit_id + access_token). These are non-obvious, high-value behavioral facts an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and mechanism, then layers cost, constraints, security, and return values. Every sentence carries information, though the description is dense and packs several distinct concerns without visual separation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a cost-incurring creation tool with no output schema and a nested weights object, the description covers the essential gaps: what it produces, cost/failure behavior, security handling, and the required follow-up call. Minor omissions remain around the battery-selection and teleology-name parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 parameters, so the description must compensate. It explains the critical constraint on teleology_weights ('must sum to 1.0 and exactly match the battery's objectives') and clarifies api_key and target_auth_token handling, but leaves target_endpoint, teleology_name, teleology_statement, and the battery_id default of 'core-v1' unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a psyche audit against a target agent endpoint') and then explains the mechanism — dilemma battery, self-reports, three-source mirror. This lets an agent distinguish it from remirror_get_report (polling) and remirror_list_batteries without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: it costs $79.00 in prepaid credits, failed audits are refunded, and the workflow routes onward explicitly ('poll remirror_get_report for the report'). It does not state when this tool should be avoided or how it relates to remirror_list_batteries, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remirror_get_reportBInspect
Fetch the status or final mirror report of an audit session. Needs the audit_id and the access_token returned at creation.
| Name | Required | Description | Default |
|---|---|---|---|
| audit_id | Yes | ||
| access_token | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose one useful fact beyond the schema — the access_token comes from creation — but says nothing about whether this is read-only, whether the report may be unavailable/partial when polled, or any rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the core purpose is front-loaded ahead of the parameter note. Nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description should carry more of the return-value story than it does — it never says what a 'status' versus a 'final report' response looks like. For a simple two-param fetch tool it is adequate but leaves a clear gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and both parameters are bare strings, so the description must compensate. It adds real meaning for access_token by stating it is 'returned at creation,' but gives no elaboration on audit_id beyond its name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (the status or final mirror report of an audit session), which is distinct enough from sibling remirror_create_audit. It does not explicitly name a sibling to differentiate against, but the resource itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase 'status or final report' hints at a lifecycle where status is polled before the final report exists, and the token requirement implies it follows remirror_create_audit. There is no explicit when-to-use, when-not-to-use, or alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remirror_list_batteriesAInspect
List available dilemma batteries for psyche audits: id, version, objectives covered, dilemma count. Free.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses that the operation is free and lists the return fields, but does not describe authentication requirements, rate limits, or whether the list is paginated or filtered. Adequate for a simple read-only list but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the action and resource, then the returned fields and cost. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description enumerates the key return fields (id, version, objectives covered, dilemma count) and notes the free cost, which compensates for the missing schema. For a zero-parameter list tool, this is nearly complete; only granular details like sorting are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter syntax to explain; the baseline is 4. The description adds no parameter meaning beyond noting the returned fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('dilemma batteries for psyche audits'), and enumerates returned fields (id, version, objectives covered, dilemma count). Sibling tools remirror_create_audit and remirror_get_report are clearly different operations, so an agent can distinguish this catalog call without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides context ('for psyche audits') and a cost note ('Free'), implying use before audit creation, but does not explicitly say when to choose this over siblings or when not to use it. With zero parameters, usage is largely inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
remirror_create_audit - First observed
remirror_get_report - First observed
remirror_list_batteries
Related MCP Connectors
Testing, benchmarking and auditing autonomous AI agents — methods, harnesses, evidence
Decision-assurance for AI agents: an auditable action boundary + receipt before it acts.
Trust gate for AI agents: multi-model adversarial consensus, signed and verifiable verdicts.
Bitcoin-anchored, tamper-evident audit log for AI agents — record, disclose and verify actions.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceLogs AI agent actions, computes reliability scores, generates audit reports, detects anomalies, and estimates risk to ensure agent accountability.-
- AlicenseNot gradedqualityCmaintenanceA human-in-the-loop governance interlock for AI agents. Agents propose changes, a human countersigns the exact plan, and then it executes stage by stage with precondition checks, verification, and auditing.Apache 2.0
- AlicenseAqualityAmaintenanceAutonomous cross-source inconsistency and contradiction intelligence engine for AI agents. Detects specification drift, SemVer conflicts, and infrastructure mismatches across code, docs, and deployments.14125MIT
- AlicenseAqualityBmaintenanceAI agent provenance, trust, and auditability layer. VERITAS multi-gate scoring, Cortex approval gates, S.E.A.L. hash-chain audit ledger, and semantic RAG with cryptographic provenance tracking for every decision an agent makes.275MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.