cross-review
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| GROK_API_KEY | Yes | xAI Grok API key | |
| GEMINI_API_KEY | Yes | Google Gemini API key | |
| OPENAI_API_KEY | Yes | OpenAI API key for Codex/OpenAI peer | |
| DEEPSEEK_API_KEY | Yes | DeepSeek API key | |
| ANTHROPIC_API_KEY | Yes | Anthropic API key for Claude | |
| CROSS_REVIEW_STUB | No | Set to '1' for stub mode (smoke tests, no cost) | 0 |
| PERPLEXITY_API_KEY | Yes | Perplexity API key | |
| CROSS_REVIEW_MAX_SESSION_COST_USD | No | Maximum session cost in USD for paid calls | |
| CROSS_REVIEW_UNTIL_STOPPED_MAX_COST_USD | No | Until-stopped max cost in USD for paid calls | |
| CROSS_REVIEW_PREFLIGHT_MAX_ROUND_COST_USD | No | Preflight max round cost in USD for paid calls |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| server_infoB | Return runtime information for the API-only Cross Review MCP server, including version, data directory and active security mode. |
| runtime_capabilitiesB | Return the stable cross-review runtime capability contract and active tool list. |
| probe_peersA | Query official provider APIs to discover available models for the current API keys, select the highest-capability documented model, and verify provider reachability. |
| session_initA | Create a durable cross-review session after probing provider availability and model selection. This does not call reviewer models yet. AI callers should submit raw proof through the |
| session_listA | List durable sessions saved under the local data directory. The default response is paginated and summary-only to keep stdio transports bounded; use session_read for one full session or detail='full' for a bounded page of full metadata. |
| session_readA | Read a durable session meta.json by session_id. |
| ask_peersA | Run a real API review round against selected peers. AI evidence supplied in |
| session_start_roundA | Start a real peer-review round in the background and return immediately with a session_id/job_id for polling. AI evidence supplied in |
| run_until_unanimousA | Generate or revise a draft and continue real API peer-review rounds until unanimous READY or the configured max_rounds is reached. AI evidence supplied in |
| session_start_unanimousA | Start real API generation/revision rounds in the background until unanimity, max_rounds or budget limit. AI evidence supplied in |
| session_cancel_jobA | Request cancellation for running background jobs in a durable session. The reason accepts at most 300 characters. Requires the verified capability token of the persisted session petitioner; another peer cannot cancel the job. Provider calls receive AbortSignal where the provider client supports it. |
| session_recover_interruptedA | Mark unfinished sessions with stale in-flight rounds as recovered after a MCP host restart so they can be resumed explicitly. Requires your own verified capability token, and recovers only the sessions you own. |
| session_pollA | Return durable session state and background job status without waiting for provider calls to finish. Default detail=summary keeps prior-round peer text/raw payloads out of polling responses; use detail=full or session_read only when full forensic data is required. |
| session_metricsA | Return aggregate observability metrics across all sessions, or only one session when session_id is provided. |
| session_peer_reliability_reportA | Read-only per-peer reliability telemetry: READY/NEEDS_EVIDENCE/NOT_READY counts, parser warnings, provider errors, unresolved evidence asks, fabrication events, latency and cost. Observational only; does not change peer selection or mutate sessions. |
| session_doctorA | Operational audit across durable sessions: open/stale/blocked cases, legacy self-lead metadata, open evidence asks (with per-peer item type drill-down + chronic blockers since v2.22), Grok provider errors, and token-event noise. Read-only by default (does not modify sessions). Terminal max-rounds and terminal not_resurfaced history stay in totals but are not default operational findings; pass include_terminal_findings=true to enumerate that historical inventory. Pass include_legacy=true to enumerate per-session self_lead_metadata entries (hidden by default since v2.22 because pre-v2.16 sessions carry the legacy artifact at ~38% rate; totals.self_lead_metadata count is always visible). v3.6.0: pass repair=true (opt-in) to recompute convergence_health for sessions stuck in the contradictory outcome="converged"+health="blocked" state left by pre-v3.2.0 corruption — only that specific contradiction is touched, only when explicitly requested; the |
| session_eventsA | Read a bounded page of durable session events. Token-delta telemetry is excluded by default; opt in only for streaming forensics. Continue with next_seq while has_more is true. |
| session_reportC | Generate and save a Markdown report with convergence, peer decisions, failures, costs and latest events. |
| session_check_convergenceA | Return the latest durable convergence state, health and scope for a saved session without calling providers. |
| session_preflight_checkA | Run the same enabled evidence and truthfulness gates used by a real review round, without calling providers. Peer-submitted inline/structured evidence is checked as review material and requires no separate attachment step. |
| session_truthfulness_preflight_checkC | Backward-compatible alias for session_preflight_check. Its top-level pass now reflects both enabled runtime gates, eliminating truthfulness-only false positives. |
| session_attach_evidenceA | Attach one durable evidence artifact to an existing session, out of band from a review round. Only the session's own petitioner may call it, and the artifact carries the same |
| session_evidence_judge_passA | LLM satisfied-detection for the Evidence Broker. The configured judge peer reads each currently-open checklist item against the supplied draft and returns a structured judgment; a peer can never judge its own evidence ask. The runtime promotes only items where satisfied=true AND confidence='verified'; everything else stays open. Terminal statuses and already-addressed items are never touched. Optional shadow_mode records non-mutating decisions. Requires the verified capability token of the persisted session petitioner, because the pass spends that petitioner's budget on paid provider calls. |
| session_evidence_judge_consensus_passA | Multi-peer evidence judgment. Requires at least two distinct enabled judge peers. A peer is forbidden from ruling on its own evidence ask; any self-judge member makes that item's consensus fail closed. Active mode promotes only unanimous verified-satisfied judgments with non-empty rationales and zero parser warnings; shadow mode never mutates state. Requires the verified capability token of the persisted session petitioner, because the pass spends that petitioner's budget on paid provider calls. |
| session_judgment_precision_reportA | v2.14.0 — compute precision/recall/F1 of the shadow judge against the empirical ground truth (whether peers raised the same ask in a subsequent round). Walks |
| contest_verdictA | v2.14.0 — formally contest a final verdict and open a new deliberation cycle. The reason accepts at most 4,000 characters. Requires the verified capability token of the persisted session petitioner (pass |
| session_sweepA | Finalize unfinished sessions whose metadata has been idle for at least 24 hours. The terminal reason accepts at most 200 characters. v3.7.5 (B1): opt-in |
| session_finalizeA | Close a non-terminal durable session as |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 28 tools
Most tools have distinct roles, but there is meaningful overlap among state-inspection tools (session_read, session_poll, session_events, session_check_convergence) and between the synchronous/asynchronous round-starting variants. The deliberate alias session_truthfulness_preflight_check also adds avoidable ambiguity.
The dominant session_* family is consistently snake_case and action-oriented, and even the non-session tools use clear lowercase_snake names. Minor deviations like server_info, runtime_capabilities, and run_until_unanimous break the strict pattern but are not seriously confusing.
At 28 tools, the surface is above the comfortable range and feels bloated, with several maintenance, telemetry, and inspection variants that could plausibly be consolidated. While the domain is complex, the alias tool and overlapping read/state operations suggest the count is not fully justified.
The lifecycle is well covered: session creation, reading, polling, round execution, evidence handling, judgment, convergence checking, reporting, cancellation, recovery, finalization, and contestation are all present. Minor gaps like explicit session export or deletion exist, but the append-only session design makes those omissions reasonable.