QA Orchestrator
QA Orchestrator is a deterministic FastMCP service that provides bounded, content-free orchestration for host-owned QA reviews via five MCP tools.
Retrieve fixed review checklists for any specialist profile (focus, required sections, constraints, escalation signals) with
prepare_review_route.Start a tracked QA session by task type, receiving recommended bundles, allowed profiles/bundles, and per-stage model policy with
start_qa_orchestration.Advance through triage, sequential primary review profiles, optional deep review, and final synthesis using structured signals; the orchestrator validates transitions and applies fixed escalation rules.
Read current session state (content-free) when resuming or recovering; sessions expire by TTL (default 1800s) and can optionally persist to SQLite for recovery.
Record the host’s final outcome (
completed,partial,blocked) after synthesis or an early stop withfinish_qa_orchestration.Host retains responsibility for evidence, model calls, findings, severity, release decisions, and external writes; the orchestrator only routes, validates, and stores bounded session state.
Allows configuring QA orchestration stages to use OpenAI models, with per-stage model and reasoning policies returned to the host for execution.
QA Orchestrator
QA Orchestrator is a small deterministic FastMCP service for host-owned QA reviews. It keeps orchestration state bounded; evidence, source code, logs, prompts, model responses, and final decisions remain with the primary host agent.
How it works
The primary host obtains authoritative evidence from the required systems and classifies the QA task.
The host calls
start_qa_orchestration. The orchestrator creates a content-free session and returns the first step and its configured policy. Runqa-orch setupto choose OpenAI or Anthropic and configure models and reasoning for each stage.The host runs each stage in its configured model environment and sends the orchestrator only a structured signal after each stage:
triage — select one fixed review bundle or one compatibility profile;
primary review — review every selected profile in the fixed order;
optional deep review — perform one read-only analysis when fixed risk signals match;
final synthesis — consolidate the results. Each stage uses the model and reasoning configured for it in the returned
model_policy. The returned policy selects the provider, model, and reasoning for each stage. Execution speed and latency preferences remain controlled by the user's host/provider settings; the orchestrator does not set or override them.
The host validates findings, runtime evidence, and limitations. For an orchestrated task, it calls
finish_qa_orchestrationwith the samerun_idand its final outcome.
The orchestrator does not call models, choose severity or release readiness, or perform external writes.
Related MCP server: Support Ticket Triage MCP
Choose a review path
Use one compatibility profile by default for a routine, narrowly scoped,
low-risk change with one main concern. Use a fixed bundle for broad,
cross-concern, or high-risk changes. For example, a small Ruby guard change can
use ruby_reviewer; a change spanning a Rails endpoint, a background job, and
their tests fits ruby_backend.
recommended_bundles is a task-type shortlist, not a risk score or a required
selection. The host inspects the diff and chooses the initial profile or
bundle. The host sends structured risk signals after primary review; the
orchestrator applies its fixed rules to decide whether optional deep review
follows.
Example: the same review with and without the orchestrator
Ordinary review guided by | Review with QA Orchestrator | |
Small Ruby guard change | The host chooses a reviewer prompt and tracks the review in the conversation. | The host selects |
Broad Ruby change across an endpoint, job, and tests | The host coordinates review steps from its instructions. | The host selects |
Deeper review | The host decides from its own instructions and evidence. | The host sends structured risk signals; the orchestrator applies fixed escalation rules. |
Ownership | The host gathers evidence and makes the final decision. | The host still owns evidence, model calls, findings, and the final decision; the orchestrator receives no source or raw evidence. |
The orchestrator adds a validated workflow contract and bounded progress state. It does not replace host instructions or perform the review itself.
QA Orchestrator at a glance

Bundles and profile names
The triage stage selects one fixed bundle or one compatibility profile. start_qa_orchestration returns recommended_bundles as a task-type-based shortlist; it does not restrict allowed_bundles. Choose the route from the changed files and confirmed project stack. The primary-review stage executes bundle profiles sequentially. After every role, the host sends completed_profile, and the orchestrator returns current_profile and completed_profiles. Use the updated session returned by advance_qa_orchestration for the next action and model policy; call get_qa_orchestration only when resuming or recovering a session. Deep review or synthesis is available only after the final role. Transition identifiers are model-neutral; choose the model from the returned model_policy, never from the step name.
Bundle | Profile order |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
autotest and widget include a TypeScript review role; select them when TypeScript review is relevant. Use ordinary_mr for broad non-TypeScript automation, widget_js for broad JavaScript React changes, ruby_backend for broad Ruby backend changes, python_backend for broad Python/MCP changes, and mobile for broad React Native or native iOS/Android changes. Choose from the changed files and confirmed project manifests, not the repository name alone; monorepos can contain several stacks.
Technical profiles and display names:
Profile | Host-facing name |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Faraday is only the internal display name of the code_explorer profile. No external agent, service, package, or model is connected under that name.
Ordinary MR flow with optional escalation:
Triage → Ordinary MR Review
Primary review → Faraday — Evidence Investigator → Code Reviewer → Test Analyzer
├─ no escalation ───────────────────────────────→ Final synthesis
└─ fixed risk signals match → Deep review → Final synthesis
Host → Final QA outcomeWith the final primary-review role, the host must send the structured boolean risk_signals object in the same advance_qa_orchestration call as the final completed_profile (send {} when no signals apply). Deep review is derived only from the fixed signal rules. Do not send risk_signals on the later synthesis transition. The returned session includes matched rules and fixed reason codes as deep_assessment; raw evidence never enters the orchestrator.
Deep-review rules:
high_risk_domain+evidence_uncertain;any two of
cross_system_scope,multiple_plausible_causes,non_reproducible, andhigh_blast_radius;evidence_conflicttogether withhigh_risk_domain,cross_system_scope, orhigh_blast_radius.
For low-risk, narrow reviews, the triage stage may select one compatibility profile instead of a bundle: code_reviewer for a small behavior change, pr_test_analyzer for a test-only change, typescript_reviewer for a TypeScript-only change, react_reviewer for a React-only change, ruby_reviewer for a Ruby-only change, python_reviewer for a Python/MCP-only change, or mobile_reviewer for a React Native/native-platform-only change. Broad or cross-concern reviews continue to use a fixed bundle.
Keep one compact per-task Evidence Packet with stable evidence references (E1, E2, ...) and bounded finding candidates (F-01, F-02, ...). Do not repeat the full diff or raw logs in every model stage.
MCP interface
The service publishes exactly five tools:
Tool | Purpose |
| Return the fixed checklist for one profile without creating a session |
| Create a session and return the task-based bundle shortlist |
| Complete triage, primary review, deep review, or synthesis and return the next action |
| Read the current content-free state when resuming or recovering a session |
| Record the host-owned outcome after synthesis or an early stop |
Bundle orchestration flow:
Triage → Primary review[1] → ... → Primary review[N]
↘ optional Deep review ↗
Final synthesis → Host outcomeSessions are kept in process memory by default. Successful state changes refresh the 1,800-second default TTL; reads do not. The maximum is 100 active sessions, and the shared cache is bounded, so older terminal sessions may be evicted when capacity is needed. Repeating the final call is idempotent while its session is retained. After a restart, the host starts a new session unless optional recovery storage is enabled. read_only=true and host_owns_decisions=true are part of every state.
Set QA_ORCHESTRATOR_SESSION_STORE_PATH to opt into a local SQLite file that restores unfinished orchestration state after a restart. It stores only the current structured session needed for recovery; finalization removes that row. It never stores evidence, prompts, source, logs, model responses, finalized outcomes, history, or statistics. Use one server process per store file. The default remains memory-only.
After synthesis, the session waits for the host's final outcome. Call finish_qa_orchestration with the session run_id and completed, partial, or blocked. For an early stop, first pass partial or blocked to advance_qa_orchestration, then finish the session with the same outcome. Repeating the same finalization is idempotent; a conflicting outcome is rejected. Tasks that do not use orchestration need no finalization call. The service does not store or report task statistics.
Review profiles
prepare_qa_orchestration returns the focus, required sections, constraints, escalation signals, and display name for one profile. For a bundle, the host calls the route for every profile in the returned fixed order and passes its technical identifier in completed_profile after each call.
required_sections is profile-specific: the evidence investigator returns Scope, Evidence Map, and Unverified; test analysis returns Scope, Coverage Gaps, and Unverified; implementation and specialist reviews return Scope, Finding Candidates, Coverage Gaps, and Unverified.
Use the local profile evaluation pack to smoke-check role boundaries and bundle selection without collecting task statistics.
Responsibility boundary
The primary host is responsible for:
obtaining and validating evidence;
calling the issue tracker, code host, test-management system, observability and logging systems, documentation and chat systems, code index, and the file system;
running triage, primary-review, synthesis, and any optional deep-review stage under the policy;
confirmed findings, severity, release/readiness judgment, and the final QA response;
file changes and all external writes.
QA Orchestrator is responsible only for fixed routing, state transitions, read-only constraints, and bounded session finalization. advance_qa_orchestration must not receive an Evidence Packet, prompt, model output, source text, logs, paths, or an arbitrary reason.
Requirements
Python 3.12+;
uv;an MCP client that supports STDIO;
Windows, macOS, or Linux.
On Windows, uv sync installs the native MCP entry point at
.venv\Scripts\qa-orchestrator-mcp.exe. macOS and Linux use the POSIX
source launcher.
Installation
For the complete setup—including Codex and Claude Code registration, host instructions, verification, and troubleshooting—see the installation guide.
git clone https://github.com/Rbkmen/qa-orchestrator.git
cd qa-orchestrator
uv sync
uv run qa-orchestrator-doctor
uv run qa-orch setupOn macOS and Linux, the source launcher uses the project's .venv, the active
VIRTUAL_ENV, or an installed qa-orchestrator-mcp from PATH. On Windows,
use the installed .venv\Scripts\qa-orchestrator-mcp.exe entry point. No
separate background process is required.
For clients that support uvx, a checkout is optional:
Replace <commit-sha> with the full commit hash and use the same hash in the
setup and server commands.
uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' qa-orch setupThis saves the model policy in your user configuration. Register the server command with your MCP client so the client starts it when needed; do not run the server command directly in a terminal. For Codex without a checkout:
codex mcp add qa-orchestrator -- uvx \
--from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' \
qa-orchestrator-mcpUse an immutable commit SHA instead of the default branch for reproducible team configuration. See the installation guide for the remaining no-checkout commands and update instructions.
The setup wizard first lets you choose Russian or English for that run, then
stores only the provider label, model IDs, and the selected provider-specific
reasoning/effort values in the local model-policy.json; the language is not
saved, and the wizard never asks for or stores API keys. Choose a provider and
model that your host client can use; the wizard records the policy but does not
configure provider access. From a checkout, use uv run qa-orch config show and
uv run qa-orch reload. Without a checkout, prefix those commands with
uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>'.
Colors in the wizard distinguish providers, model IDs, and reasoning values;
set NO_COLOR=1 to disable them or FORCE_COLOR=1 to force them.
The wizard uses a local recommendation/capability catalog and does not contact provider APIs. For custom model IDs, verify that the host account can access the model and supports the selected reasoning/effort value.
Example for Codex with a local checkout:
codex mcp add qa-orchestrator -- "$(pwd)/scripts/qa-orchestrator"On Windows, run this from the repository root in PowerShell:
codex mcp add qa-orchestrator -- "$PWD\.venv\Scripts\qa-orchestrator-mcp.exe"The two supported host integrations are described in the client guides.
Configuration
Variable | Default |
|
|
|
|
|
|
|
|
| unset (disabled) |
Finalize an orchestration
For completion and early-stop instructions, see the MCP interface section.
Clients and rules
Development
See CONTRIBUTING.md. Any change to the public MCP contract must include an exact tool-surface test and a check that evidence, decisions, and external writes remain with the primary host.
Available Tools
5 toolsadvance_qa_orchestrationA
Complete the active review step or mark an early stop; return the next action.
Pass current_step as completed_step; stale or out-of-order steps are rejected. Malformed arguments, incompatible signals, and unknown or expired run_ids return errors without advancing the run. This call is non-idempotent: if a successful response is lost, inspect get_qa_orchestration before continuing; replaying the prior step is rejected.
Completed triage requires exactly one of selected_bundle or selected_profile; omit both when stopping early.
After each completed primary review, send current_profile as completed_profile. On the final profile, also send risk_signals ({} if none apply); omit both on an early stop.
An early stop records partial or blocked but still needs finish_qa_orchestration with the same outcome. After synthesis, use finish_qa_orchestration for the host's final outcome.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it. | |
| status | Yes | Use completed after the active step succeeds. Use partial or blocked to stop early; then finalize with the same outcome using finish_qa_orchestration. | |
| risk_signals | No | Required with the final completed primary-review profile. Send an empty object when no evidence-based fixed escalation signals apply; do not send on earlier profiles or on the synthesis transition. | |
| completed_step | Yes | Active step returned in current_step and completed by this call: triage, primary_review, deep_review, or synthesis. When current_step is awaiting_host_outcome, call finish_qa_orchestration instead. | |
| selected_bundle | No | For completed triage, choose one allowed fixed bundle for a broad or cross-concern review. It is mutually exclusive with selected_profile; recommended_bundles is only a shortlist. | |
| selected_profile | No | For completed triage, choose one allowed profile for a narrow, low-risk review. It is mutually exclusive with selected_bundle. | |
| completed_profile | No | After each successful primary-review step, submit exactly the current_profile returned by the previous advance_qa_orchestration call. Omit for other steps or an early stop. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-idempotency (idempotentHint=false), but the description goes further by spelling out the consequence: a lost successful response must be recovered via get_qa_orchestration because replaying the prior step is rejected. It also discloses validation behavior (malformed args, incompatible signals, unknown/expired run_ids error without advancing) and state mutation (early stop records partial/blocked), all consistent with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The summary is front-loaded in the first sentence, followed by an ordered bullet list of transition rules. It is dense but each bullet carries a distinct rule; sizing is slightly heavy but justified by the workflow complexity, with no pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return shape needn't be described, and the description covers the remaining gaps: step sequencing, mutual exclusivity, escalation-signal timing, and recovery after a lost response. Nothing an agent needs to drive this multi-step orchestration correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even with 100% schema coverage, the description adds cross-parameter conditional logic the schema cannot express: exactly one of selected_bundle or selected_profile on completed triage (omit both on early stop), completed_profile on every primary-review step, and risk_signals only on the final profile with {} when none apply. This sequencing semantics materially exceeds the per-field schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a precise verb+resource ('Complete the active review step or mark an early stop; return the next action') and immediately clarifies it is the step-advancing tool. It explicitly contrasts with siblings by naming finish_qa_orchestration and get_qa_orchestration and the conditions under which each applies, so an agent can disambiguate without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states explicit routing rules: use finish_qa_orchestration when current_step is awaiting_host_outcome, and use it again for the host's final outcome after synthesis. It also covers when-not scenarios (stale or out-of-order steps rejected, early stop still requires finish_qa_orchestration) with concrete alternative conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
finish_qa_orchestrationAIdempotent
Finalize the host's outcome after synthesis or an early stop; no review step runs here.
Call when advance_qa_orchestration reaches awaiting_host_outcome after synthesis, or after
it records an early stop. For an early stop, submit the same partial or blocked outcome.
Repeating the same outcome returns the retained terminal session. A conflicting outcome is rejected with
conflicting final outcomeand leaves the stored status unchanged.Retention ends when the configured TTL expires (1800 seconds by default), a terminal session is evicted to free capacity, or the service restarts. After removal, the run_id is unavailable and the outcome can no longer be deduplicated.
This tool records the host's decision; it does not make it.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it. | |
| outcome | Yes | Host-selected final session status: completed, partial, or blocked. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare idempotence and non-destructiveness, and the description goes well beyond them: repeat outcomes return the retained terminal session, conflicts are rejected with a specific error and leave the stored status unchanged, and retention ends on TTL expiry (1800s), eviction, or restart. It also disclaims agency ('it does not make it'). This is unusually rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then conditions, then behavioral rules as bullets. Every sentence carries non-redundant information; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not needed. The description covers the lifecycle (when to call), error semantics, idempotence/dedup, and retention limits — everything needed to invoke and reason about the call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters are documented in the schema (run_id pattern, outcome enum). The description still adds meaning beyond the schema by prescribing which outcome values are valid for the early-stop path, which is value-selection guidance not present in the field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Finalize the host's outcome after synthesis or an early stop.' It also disambiguates against the sibling advance_qa_orchestration by denying scope ('no review step runs here'), so an agent can separate it from the other orchestration tools without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger conditions are given: call when advance_qa_orchestration reaches `awaiting_host_outcome` after synthesis, or after it records an early stop. It also states the correct action for the early-stop path (submit the same `partial` or `blocked` outcome). No inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_qa_orchestrationARead-onlyIdempotent
Read a session's current content-free state without changing it.
Use this only to inspect a retained session's current step, next action, and model policy. Reads do not extend the TTL; an unknown or expired session cannot be recovered here, so start a new one. Use advance_qa_orchestration to change an active run or finish_qa_orchestration to record its final outcome.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint, but the description adds real behavioral context beyond them: reads do not extend the TTL, expired or unknown sessions are unrecoverable through this tool, and the state inspected is content-free. These are non-obvious operational traits an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool does, then immediately adds the critical constraint (no TTL extension, no recovery) and routing to alternatives. Three tight sentences, zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation; the description instead covers the runtime semantics an agent can't get from structured fields (TTL behavior, expired-session failure mode, when to use siblings). Complete for a single-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the run_id pattern/format is fully documented in the schema itself, so the description adds nothing about the parameter. Baseline 3 is appropriate when the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read a session's current content-free state'), plus the exact fields returned (current step, next action, model policy). It is immediately distinguishable from siblings named in the text (advance_qa_orchestration, finish_qa_orchestration).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes usage ('only to inspect a retained session'), names the two alternatives and what each is for, and states the negative case ('an unknown or expired session cannot be recovered here, so start a new one'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_qa_orchestrationARead-onlyIdempotent
Return a fixed checklist for one agent_profile. It gives review focus, required output
sections, shared evidence constraints, and escalation signals.
Use it for one scoped concern. For a selected bundle, request each review_profiles member
separately in session order; use start_qa_orchestration to track the multi-concern review.
Passing agent_profile never selects a session profile. This stateless lookup does not inspect
repository content, run the review, or change session state; the host owns final decisions.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_profile | Yes | Choose one fixed specialist profile: code_explorer maps execution and evidence; code_reviewer reviews implementation; pr_test_analyzer reviews coverage; security_reviewer and silent_failure_hunter target security and silent failures; ruby_reviewer, python_reviewer, typescript_reviewer, react_reviewer, and mobile_reviewer target their named stacks. Choose from changed files and confirmed stack. |
Output Schema
| Name | Required | Description |
|---|---|---|
| focus | Yes | The profile's review boundary and primary responsibility. |
| profile | Yes | |
| read_only | No | |
| constraints | Yes | Shared limits for this read-only review profile. |
| display_name | Yes | |
| required_sections | Yes | Profile-specific output sections; use these instead of a generic template. |
| escalation_signals | No | Conditions to assess against evidence. A listed condition is not itself a deep-review trigger; set risk_signals only when the corresponding evidence-based condition applies. |
| host_owns_decisions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/closed-world annotations, the description discloses that this is a stateless lookup that does not inspect repository content, run the review, or change session state, and that the host owns final decisions. It also pre-empts a likely misreading: passing agent_profile never selects a session profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core behavior in the first sentence, then usage guidance, then scope constraints. Slightly dense and broken across lines, but every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, yet the description still characterizes the checklist's contents. Combined with the statelessness and session-selection caveats, an agent has everything needed to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum description already maps each profile to its concern, so the baseline is 3. The description adds real semantic value by clarifying what agent_profile does not do (it never selects a session profile), which is not derivable from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return a fixed checklist for one agent_profile') and enumerates what the checklist contains (review focus, required output sections, shared evidence constraints, escalation signals). It is clearly distinguishable from the session-lifecycle siblings (start/advance/finish/get_qa_orchestration).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes usage to 'one scoped concern' and names the alternative path: request each review_profiles member separately in session order for a bundle, and use start_qa_orchestration to track the multi-concern review. Both when-to-use and the sibling alternative are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_qa_orchestrationA
Create a content-free QA session for a tracked review; call once before triage.
Every call creates a separate session. Use prepare_qa_orchestration for a stateless profile
checklist. task_type narrows the bundle shortlist; the host still selects the review route.
Sessions expire after the configured TTL (1800 seconds by default).
Capacity is 100 sessions by default and can be configured.
Expired sessions are purged and retained terminal sessions may be evicted. If capacity can only be freed by removing an active session, the call returns a
session limiterror.
| Name | Required | Description | Default |
|---|---|---|---|
| task_type | Yes | Broad request category: ordinary_review for implementation review, widget_review for widget changes, epic_analysis for broad exploration, requirements_analysis for requirement review, qa_planning for coverage planning, autotest_implementation for automation work, or other. This affects recommended_bundles only; the host chooses the route from changed files and evidence; it is not a risk level. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover safety/idempotency flags; the description adds substantial behavioral context beyond them: TTL expiry (1800s default), 100-session capacity, purge/eviction behavior, and the specific 'session limit' error when only an active session could free capacity. This is exactly the operational detail annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then a routing sentence, then tight bulleted constraints. Every sentence carries distinct information with no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema covers return values, so the description is not obliged to describe them. Combined with annotations and the schema, the description supplies the lifecycle (TTL, capacity, eviction, error) an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum is fully self-documented, so baseline 3 applies. The description still adds real meaning by clarifying what task_type does and does not do ('narrows the bundle shortlist', 'host still selects the review route', 'not a risk level'), correcting a plausible misreading.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (content-free QA session) with the timing constraint (call once before triage) front-loaded. It explicitly contrasts with prepare_qa_orchestration, so the agent can separate it from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when ('call once before triage') and when-to-use-the-alternative ('Use prepare_qa_orchestration for a stateless profile checklist'). It also clarifies that every call creates a separate session, preempting accidental reuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.12- Added
prepare_qa_orchestration - Removed
prepare_review_route
1 tool update
v0.1.6- Changed
finish_qa_orchestration1 field changed- changed
Input schema / properties / outcome / descriptionPrevious value: -"Host-selected final QA outcome: completed after synthesis, or the same partial or blocked outcome already used for an early stop. Conflicting final outcomes are rejected."New value: +"Host-selected final session status: completed, partial, or blocked."
5 tool updates
v0.1.1- Changed
advance_qa_orchestration8 fields changed- added
Input schema / properties / completed_profile / descriptionAdded value: +"After each successful primary-review step, submit exactly the current_profile returned by the previous advance_qa_orchestration call. Omit for other steps or an early stop." - added
Input schema / properties / completed_step / descriptionAdded value: +"Active step returned in current_step and completed by this call: triage, primary_review, deep_review, or synthesis. When current_step is awaiting_host_outcome, call finish_qa_orchestration instead." - changed
Input schema / properties / completed_step / enumPrevious value: -[ - "triage", - "primary_review", - "deep_review", - "synthesis", - "awaiting_host_outcome" -]New value: +[ + "triage", + "primary_review", + "deep_review", + "synthesis" +] - changed
Input schema / properties / risk_signals / descriptionPrevious value: -"Required with the final completed primary-review profile. Send an empty object when no signals apply."New value: +"Required with the final completed primary-review profile. Send an empty object when no evidence-based fixed escalation signals apply; do not send on earlier profiles or on the synthesis transition." - added
Input schema / properties / run_id / descriptionAdded value: +"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it." - added
Input schema / properties / selected_bundle / descriptionAdded value: +"For completed triage, choose one allowed fixed bundle for a broad or cross-concern review. It is mutually exclusive with selected_profile; recommended_bundles is only a shortlist." - added
Input schema / properties / selected_profile / descriptionAdded value: +"For completed triage, choose one allowed profile for a narrow, low-risk review. It is mutually exclusive with selected_bundle." - added
Input schema / properties / status / descriptionAdded value: +"Use completed after the active step succeeds. Use partial or blocked to stop early; then finalize with the same outcome using finish_qa_orchestration."
- Changed
finish_qa_orchestration2 fields changed- added
Input schema / properties / outcome / descriptionAdded value: +"Host-selected final QA outcome: completed after synthesis, or the same partial or blocked outcome already used for an early stop. Conflicting final outcomes are rejected." - added
Input schema / properties / run_id / descriptionAdded value: +"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
- Changed
get_qa_orchestration1 field changed- added
Input schema / properties / run_id / descriptionAdded value: +"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
- Changed
prepare_review_route1 field changed- added
Input schema / properties / agent_profile / descriptionAdded value: +"Choose one fixed specialist profile: code_explorer maps execution and evidence; code_reviewer reviews implementation; pr_test_analyzer reviews coverage; security_reviewer and silent_failure_hunter target security and silent failures; ruby_reviewer, python_reviewer, typescript_reviewer, react_reviewer, and mobile_reviewer target their named stacks. Choose from changed files and confirmed stack."
- Changed
start_qa_orchestration1 field changed- added
Input schema / properties / task_type / descriptionAdded value: +"Broad request category: ordinary_review for implementation review, widget_review for widget changes, epic_analysis for broad exploration, requirements_analysis for requirement review, qa_planning for coverage planning, autotest_implementation for automation work, or other. This affects recommended_bundles only; the host chooses the route from changed files and evidence; it is not a risk level."
5 tool updates
v0.1.0- First observed
advance_qa_orchestration - First observed
finish_qa_orchestration - First observed
get_qa_orchestration - First observed
prepare_review_route - First observed
start_qa_orchestration
TDQS
Scored across 5 tools
The five tools map to distinct lifecycle stages (prepare=stateless lookup, start=create session, advance=progress step, finish=terminal outcome, get=read state). The main risk is confusing prepare_qa_orchestration with start_qa_orchestration, but the descriptions explicitly disambiguate by stressing prepare is stateless and never creates a session. advance vs finish is also clarified by tying finish to the awaiting_host_outcome terminal state.
All five tools follow a strict verb_qa_orchestration pattern: start_, prepare_, advance_, finish_, get_. The convention is uniform across the entire set with no deviations or mixed casing.
Five tools is a tight, well-scoped set that exactly covers a session lifecycle plus a stateless helper. Each tool earns its place with no redundancy or padding.
The surface covers the full lifecycle (create, stateless lookup, progress, read, terminal finalize) and handles early stops. Minor gap: there is no explicit list/abort/cancel operation, though early stop is folded into advance + finish, so agents can still work around it.
Maintenance
Related MCP Connectors
Hosted MCP for denial, prior auth, reimbursement, workflow validation, batch scoring, and feedback.
The MCP server that vets MCP servers: identity, risk grade and per-tool risk before you install.
Hybrid human + AI expertise for faster, trusted answers and decisions via MCP Server.
A paid remote MCP for hosted MCP server, built to return verdicts, receipts, usage logs, and audit-r
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server for healthcare claims workflow scoring, validation, and feedback, supporting denial risk, prior authorization, and reimbursement assessment.838 PyPIMIT
- FlicenseAqualityFmaintenanceA local MCP server for governed support-ticket triage that reads synthetic tickets and knowledge articles, prepares evidence-backed recommendations, and records local audit events.9-
- AlicenseCqualityDmaintenanceAn MCP server that guides QA and verification processes by breaking down tasks into manageable steps and providing LLM-driven, confidence-scored tool recommendations.112 npm6MIT
- AlicenseAqualityCmaintenanceA self-hosted MCP server offering provider-independent AI routing, deterministic verification, and cryptographically signed, tamper-evident audit trails for every response.831 npm2MIT