Skip to main content
Glama

QA Orchestrator

QA Orchestrator MCP server – quality and maintenance score on Glama

QA Orchestrator is a small deterministic FastMCP service for host-owned QA reviews. It keeps orchestration state bounded; evidence, source code, logs, prompts, model responses, and final decisions remain with the primary host agent.

How it works

  1. The primary host obtains authoritative evidence from the required systems and classifies the QA task.

  2. The host calls start_qa_orchestration. The orchestrator creates a content-free session and returns the first step and its configured policy. Run qa-orch setup to choose OpenAI or Anthropic and configure models and reasoning for each stage.

  3. The host runs each stage in its configured model environment and sends the orchestrator only a structured signal after each stage:

    • triage — select one fixed review bundle or one compatibility profile;

    • primary review — review every selected profile in the fixed order;

    • optional deep review — perform one read-only analysis when fixed risk signals match;

    • final synthesis — consolidate the results. Each stage uses the model and reasoning configured for it in the returned model_policy. The returned policy selects the provider, model, and reasoning for each stage. Execution speed and latency preferences remain controlled by the user's host/provider settings; the orchestrator does not set or override them.

  4. The host validates findings, runtime evidence, and limitations. For an orchestrated task, it calls finish_qa_orchestration with the same run_id and its final outcome.

The orchestrator does not call models, choose severity or release readiness, or perform external writes.

Related MCP server: Support Ticket Triage MCP

Choose a review path

Use one compatibility profile by default for a routine, narrowly scoped, low-risk change with one main concern. Use a fixed bundle for broad, cross-concern, or high-risk changes. For example, a small Ruby guard change can use ruby_reviewer; a change spanning a Rails endpoint, a background job, and their tests fits ruby_backend.

recommended_bundles is a task-type shortlist, not a risk score or a required selection. The host inspects the diff and chooses the initial profile or bundle. The host sends structured risk signals after primary review; the orchestrator applies its fixed rules to decide whether optional deep review follows.

Example: the same review with and without the orchestrator

Ordinary review guided by AGENTS.md

Review with QA Orchestrator

Small Ruby guard change

The host chooses a reviewer prompt and tracks the review in the conversation.

The host selects ruby_reviewer; the MCP validates the transition and returns the bounded session state.

Broad Ruby change across an endpoint, job, and tests

The host coordinates review steps from its instructions.

The host selects ruby_backend; the fixed reviewer order and required transitions are explicit.

Deeper review

The host decides from its own instructions and evidence.

The host sends structured risk signals; the orchestrator applies fixed escalation rules.

Ownership

The host gathers evidence and makes the final decision.

The host still owns evidence, model calls, findings, and the final decision; the orchestrator receives no source or raw evidence.

The orchestrator adds a validated workflow contract and bounded progress state. It does not replace host instructions or perform the review itself.

QA Orchestrator at a glance

Detailed host-led QA Orchestrator workflow, including review stages, optional escalation, and ownership boundaries.

Bundles and profile names

The triage stage selects one fixed bundle or one compatibility profile. start_qa_orchestration returns recommended_bundles as a task-type-based shortlist; it does not restrict allowed_bundles. Choose the route from the changed files and confirmed project stack. The primary-review stage executes bundle profiles sequentially. After every role, the host sends completed_profile, and the orchestrator returns current_profile and completed_profiles. Use the updated session returned by advance_qa_orchestration for the next action and model policy; call get_qa_orchestration only when resuming or recovering a session. Deep review or synthesis is available only after the final role. Transition identifiers are model-neutral; choose the model from the returned model_policy, never from the step name.

Bundle

Profile order

ordinary_mr

code_explorer → code_reviewer → pr_test_analyzer

widget

code_explorer → react_reviewer → typescript_reviewer → pr_test_analyzer

widget_js

code_explorer → code_reviewer → react_reviewer → pr_test_analyzer

ruby_backend

code_explorer → ruby_reviewer → pr_test_analyzer

python_backend

code_explorer → python_reviewer → pr_test_analyzer

mobile

code_explorer → mobile_reviewer → pr_test_analyzer

security

code_explorer → security_reviewer → silent_failure_hunter

autotest

code_reviewer → pr_test_analyzer → typescript_reviewer

requirements

code_explorer → code_reviewer

autotest and widget include a TypeScript review role; select them when TypeScript review is relevant. Use ordinary_mr for broad non-TypeScript automation, widget_js for broad JavaScript React changes, ruby_backend for broad Ruby backend changes, python_backend for broad Python/MCP changes, and mobile for broad React Native or native iOS/Android changes. Choose from the changed files and confirmed project manifests, not the repository name alone; monorepos can contain several stacks.

Technical profiles and display names:

Profile

Host-facing name

code_explorer

Faraday — Evidence Investigator

code_reviewer

Code Reviewer

pr_test_analyzer

Test Analyzer

security_reviewer

Security Reviewer

silent_failure_hunter

Silent Failure Hunter

typescript_reviewer

TypeScript Reviewer

react_reviewer

React Reviewer

ruby_reviewer

Ruby Reviewer

python_reviewer

Python Reviewer

mobile_reviewer

Mobile Reviewer

Faraday is only the internal display name of the code_explorer profile. No external agent, service, package, or model is connected under that name.

Ordinary MR flow with optional escalation:

Triage → Ordinary MR Review
Primary review → Faraday — Evidence Investigator → Code Reviewer → Test Analyzer
  ├─ no escalation ───────────────────────────────→ Final synthesis
  └─ fixed risk signals match → Deep review → Final synthesis
Host → Final QA outcome

With the final primary-review role, the host must send the structured boolean risk_signals object in the same advance_qa_orchestration call as the final completed_profile (send {} when no signals apply). Deep review is derived only from the fixed signal rules. Do not send risk_signals on the later synthesis transition. The returned session includes matched rules and fixed reason codes as deep_assessment; raw evidence never enters the orchestrator.

Deep-review rules:

  1. high_risk_domain + evidence_uncertain;

  2. any two of cross_system_scope, multiple_plausible_causes, non_reproducible, and high_blast_radius;

  3. evidence_conflict together with high_risk_domain, cross_system_scope, or high_blast_radius.

For low-risk, narrow reviews, the triage stage may select one compatibility profile instead of a bundle: code_reviewer for a small behavior change, pr_test_analyzer for a test-only change, typescript_reviewer for a TypeScript-only change, react_reviewer for a React-only change, ruby_reviewer for a Ruby-only change, python_reviewer for a Python/MCP-only change, or mobile_reviewer for a React Native/native-platform-only change. Broad or cross-concern reviews continue to use a fixed bundle.

Keep one compact per-task Evidence Packet with stable evidence references (E1, E2, ...) and bounded finding candidates (F-01, F-02, ...). Do not repeat the full diff or raw logs in every model stage.

MCP interface

The service publishes exactly five tools:

Tool

Purpose

prepare_qa_orchestration(agent_profile)

Return the fixed checklist for one profile without creating a session

start_qa_orchestration(task_type)

Create a session and return the task-based bundle shortlist

advance_qa_orchestration(...)

Complete triage, primary review, deep review, or synthesis and return the next action

get_qa_orchestration(run_id)

Read the current content-free state when resuming or recovering a session

finish_qa_orchestration(run_id, outcome)

Record the host-owned outcome after synthesis or an early stop

Bundle orchestration flow:

Triage → Primary review[1] → ... → Primary review[N]
                                      ↘ optional Deep review ↗
                                           Final synthesis → Host outcome

Sessions are kept in process memory by default. Successful state changes refresh the 1,800-second default TTL; reads do not. The maximum is 100 active sessions, and the shared cache is bounded, so older terminal sessions may be evicted when capacity is needed. Repeating the final call is idempotent while its session is retained. After a restart, the host starts a new session unless optional recovery storage is enabled. read_only=true and host_owns_decisions=true are part of every state.

Set QA_ORCHESTRATOR_SESSION_STORE_PATH to opt into a local SQLite file that restores unfinished orchestration state after a restart. It stores only the current structured session needed for recovery; finalization removes that row. It never stores evidence, prompts, source, logs, model responses, finalized outcomes, history, or statistics. Use one server process per store file. The default remains memory-only.

After synthesis, the session waits for the host's final outcome. Call finish_qa_orchestration with the session run_id and completed, partial, or blocked. For an early stop, first pass partial or blocked to advance_qa_orchestration, then finish the session with the same outcome. Repeating the same finalization is idempotent; a conflicting outcome is rejected. Tasks that do not use orchestration need no finalization call. The service does not store or report task statistics.

Review profiles

prepare_qa_orchestration returns the focus, required sections, constraints, escalation signals, and display name for one profile. For a bundle, the host calls the route for every profile in the returned fixed order and passes its technical identifier in completed_profile after each call.

required_sections is profile-specific: the evidence investigator returns Scope, Evidence Map, and Unverified; test analysis returns Scope, Coverage Gaps, and Unverified; implementation and specialist reviews return Scope, Finding Candidates, Coverage Gaps, and Unverified.

Use the local profile evaluation pack to smoke-check role boundaries and bundle selection without collecting task statistics.

Responsibility boundary

The primary host is responsible for:

  • obtaining and validating evidence;

  • calling the issue tracker, code host, test-management system, observability and logging systems, documentation and chat systems, code index, and the file system;

  • running triage, primary-review, synthesis, and any optional deep-review stage under the policy;

  • confirmed findings, severity, release/readiness judgment, and the final QA response;

  • file changes and all external writes.

QA Orchestrator is responsible only for fixed routing, state transitions, read-only constraints, and bounded session finalization. advance_qa_orchestration must not receive an Evidence Packet, prompt, model output, source text, logs, paths, or an arbitrary reason.

Requirements

  • Python 3.12+;

  • uv;

  • an MCP client that supports STDIO;

  • Windows, macOS, or Linux.

On Windows, uv sync installs the native MCP entry point at .venv\Scripts\qa-orchestrator-mcp.exe. macOS and Linux use the POSIX source launcher.

Installation

For the complete setup—including Codex and Claude Code registration, host instructions, verification, and troubleshooting—see the installation guide.

git clone https://github.com/Rbkmen/qa-orchestrator.git
cd qa-orchestrator
uv sync
uv run qa-orchestrator-doctor
uv run qa-orch setup

On macOS and Linux, the source launcher uses the project's .venv, the active VIRTUAL_ENV, or an installed qa-orchestrator-mcp from PATH. On Windows, use the installed .venv\Scripts\qa-orchestrator-mcp.exe entry point. No separate background process is required.

For clients that support uvx, a checkout is optional:

Replace <commit-sha> with the full commit hash and use the same hash in the setup and server commands.

uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' qa-orch setup

This saves the model policy in your user configuration. Register the server command with your MCP client so the client starts it when needed; do not run the server command directly in a terminal. For Codex without a checkout:

codex mcp add qa-orchestrator -- uvx \
  --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' \
  qa-orchestrator-mcp

Use an immutable commit SHA instead of the default branch for reproducible team configuration. See the installation guide for the remaining no-checkout commands and update instructions.

The setup wizard first lets you choose Russian or English for that run, then stores only the provider label, model IDs, and the selected provider-specific reasoning/effort values in the local model-policy.json; the language is not saved, and the wizard never asks for or stores API keys. Choose a provider and model that your host client can use; the wizard records the policy but does not configure provider access. From a checkout, use uv run qa-orch config show and uv run qa-orch reload. Without a checkout, prefix those commands with uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>'. Colors in the wizard distinguish providers, model IDs, and reasoning values; set NO_COLOR=1 to disable them or FORCE_COLOR=1 to force them.

The wizard uses a local recommendation/capability catalog and does not contact provider APIs. For custom model IDs, verify that the host account can access the model and supports the selected reasoning/effort value.

Example for Codex with a local checkout:

codex mcp add qa-orchestrator -- "$(pwd)/scripts/qa-orchestrator"

On Windows, run this from the repository root in PowerShell:

codex mcp add qa-orchestrator -- "$PWD\.venv\Scripts\qa-orchestrator-mcp.exe"

The two supported host integrations are described in the client guides.

Configuration

Variable

Default

QA_ORCHESTRATOR_DATA_DIR

~/.qa-orchestrator

QA_ORCHESTRATOR_ORCHESTRATION_TTL_SECONDS

1800

QA_ORCHESTRATOR_ORCHESTRATION_MAX_SESSIONS

100

QA_ORCHESTRATOR_MODEL_POLICY_PATH

~/.qa-orchestrator/model-policy.json

QA_ORCHESTRATOR_SESSION_STORE_PATH

unset (disabled)

Finalize an orchestration

For completion and early-stop instructions, see the MCP interface section.

Clients and rules

Development

See CONTRIBUTING.md. Any change to the public MCP contract must include an exact tool-surface test and a check that evidence, decisions, and external writes remain with the primary host.

Available Tools

5 tools
advance_qa_orchestrationA

Complete the active review step or mark an early stop; return the next action.

Pass current_step as completed_step; stale or out-of-order steps are rejected. Malformed arguments, incompatible signals, and unknown or expired run_ids return errors without advancing the run. This call is non-idempotent: if a successful response is lost, inspect get_qa_orchestration before continuing; replaying the prior step is rejected.

  • Completed triage requires exactly one of selected_bundle or selected_profile; omit both when stopping early.

  • After each completed primary review, send current_profile as completed_profile. On the final profile, also send risk_signals ({} if none apply); omit both on an early stop.

  • An early stop records partial or blocked but still needs finish_qa_orchestration with the same outcome. After synthesis, use finish_qa_orchestration for the host's final outcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesSession identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it.
statusYesUse completed after the active step succeeds. Use partial or blocked to stop early; then finalize with the same outcome using finish_qa_orchestration.
risk_signalsNoRequired with the final completed primary-review profile. Send an empty object when no evidence-based fixed escalation signals apply; do not send on earlier profiles or on the synthesis transition.
completed_stepYesActive step returned in current_step and completed by this call: triage, primary_review, deep_review, or synthesis. When current_step is awaiting_host_outcome, call finish_qa_orchestration instead.
selected_bundleNoFor completed triage, choose one allowed fixed bundle for a broad or cross-concern review. It is mutually exclusive with selected_profile; recommended_bundles is only a shortlist.
selected_profileNoFor completed triage, choose one allowed profile for a narrow, low-risk review. It is mutually exclusive with selected_bundle.
completed_profileNoAfter each successful primary-review step, submit exactly the current_profile returned by the previous advance_qa_orchestration call. Omit for other steps or an early stop.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-idempotency (idempotentHint=false), but the description goes further by spelling out the consequence: a lost successful response must be recovered via get_qa_orchestration because replaying the prior step is rejected. It also discloses validation behavior (malformed args, incompatible signals, unknown/expired run_ids error without advancing) and state mutation (early stop records partial/blocked), all consistent with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The summary is front-loaded in the first sentence, followed by an ordered bullet list of transition rules. It is dense but each bullet carries a distinct rule; sizing is slightly heavy but justified by the workflow complexity, with no pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return shape needn't be described, and the description covers the remaining gaps: step sequencing, mutual exclusivity, escalation-signal timing, and recovery after a lost response. Nothing an agent needs to drive this multi-step orchestration correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even with 100% schema coverage, the description adds cross-parameter conditional logic the schema cannot express: exactly one of selected_bundle or selected_profile on completed triage (omit both on early stop), completed_profile on every primary-review step, and risk_signals only on the final profile with {} when none apply. This sequencing semantics materially exceeds the per-field schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a precise verb+resource ('Complete the active review step or mark an early stop; return the next action') and immediately clarifies it is the step-advancing tool. It explicitly contrasts with siblings by naming finish_qa_orchestration and get_qa_orchestration and the conditions under which each applies, so an agent can disambiguate without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states explicit routing rules: use finish_qa_orchestration when current_step is awaiting_host_outcome, and use it again for the host's final outcome after synthesis. It also covers when-not scenarios (stale or out-of-order steps rejected, early stop still requires finish_qa_orchestration) with concrete alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

finish_qa_orchestrationA
Idempotent

Finalize the host's outcome after synthesis or an early stop; no review step runs here.

Call when advance_qa_orchestration reaches awaiting_host_outcome after synthesis, or after it records an early stop. For an early stop, submit the same partial or blocked outcome.

  • Repeating the same outcome returns the retained terminal session. A conflicting outcome is rejected with conflicting final outcome and leaves the stored status unchanged.

  • Retention ends when the configured TTL expires (1800 seconds by default), a terminal session is evicted to free capacity, or the service restarts. After removal, the run_id is unavailable and the outcome can no longer be deduplicated.

This tool records the host's decision; it does not make it.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesSession identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it.
outcomeYesHost-selected final session status: completed, partial, or blocked.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare idempotence and non-destructiveness, and the description goes well beyond them: repeat outcomes return the retained terminal session, conflicts are rejected with a specific error and leave the stored status unchanged, and retention ends on TTL expiry (1800s), eviction, or restart. It also disclaims agency ('it does not make it'). This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then conditions, then behavioral rules as bullets. Every sentence carries non-redundant information; nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not needed. The description covers the lifecycle (when to call), error semantics, idempotence/dedup, and retention limits — everything needed to invoke and reason about the call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are documented in the schema (run_id pattern, outcome enum). The description still adds meaning beyond the schema by prescribing which outcome values are valid for the early-stop path, which is value-selection guidance not present in the field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Finalize the host's outcome after synthesis or an early stop.' It also disambiguates against the sibling advance_qa_orchestration by denying scope ('no review step runs here'), so an agent can separate it from the other orchestration tools without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger conditions are given: call when advance_qa_orchestration reaches `awaiting_host_outcome` after synthesis, or after it records an early stop. It also states the correct action for the early-stop path (submit the same `partial` or `blocked` outcome). No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_qa_orchestrationA
Read-onlyIdempotent

Read a session's current content-free state without changing it.

Use this only to inspect a retained session's current step, next action, and model policy. Reads do not extend the TTL; an unknown or expired session cannot be recovered here, so start a new one. Use advance_qa_orchestration to change an active run or finish_qa_orchestration to record its final outcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesSession identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint, but the description adds real behavioral context beyond them: reads do not extend the TTL, expired or unknown sessions are unrecoverable through this tool, and the state inspected is content-free. These are non-obvious operational traits an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what the tool does, then immediately adds the critical constraint (no TTL extension, no recovery) and routing to alternatives. Three tight sentences, zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation; the description instead covers the runtime semantics an agent can't get from structured fields (TTL behavior, expired-session failure mode, when to use siblings). Complete for a single-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the run_id pattern/format is fully documented in the schema itself, so the description adds nothing about the parameter. Baseline 3 is appropriate when the schema does all the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read a session's current content-free state'), plus the exact fields returned (current step, next action, model policy). It is immediately distinguishable from siblings named in the text (advance_qa_orchestration, finish_qa_orchestration).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes usage ('only to inspect a retained session'), names the two alternatives and what each is for, and states the negative case ('an unknown or expired session cannot be recovered here, so start a new one'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_qa_orchestrationA
Read-onlyIdempotent

Return a fixed checklist for one agent_profile. It gives review focus, required output sections, shared evidence constraints, and escalation signals.

Use it for one scoped concern. For a selected bundle, request each review_profiles member separately in session order; use start_qa_orchestration to track the multi-concern review.

Passing agent_profile never selects a session profile. This stateless lookup does not inspect repository content, run the review, or change session state; the host owns final decisions.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_profileYesChoose one fixed specialist profile: code_explorer maps execution and evidence; code_reviewer reviews implementation; pr_test_analyzer reviews coverage; security_reviewer and silent_failure_hunter target security and silent failures; ruby_reviewer, python_reviewer, typescript_reviewer, react_reviewer, and mobile_reviewer target their named stacks. Choose from changed files and confirmed stack.

Output Schema

ParametersJSON Schema
NameRequiredDescription
focusYesThe profile's review boundary and primary responsibility.
profileYes
read_onlyNo
constraintsYesShared limits for this read-only review profile.
display_nameYes
required_sectionsYesProfile-specific output sections; use these instead of a generic template.
escalation_signalsNoConditions to assess against evidence. A listed condition is not itself a deep-review trigger; set risk_signals only when the corresponding evidence-based condition applies.
host_owns_decisionsNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent/closed-world annotations, the description discloses that this is a stateless lookup that does not inspect repository content, run the review, or change session state, and that the host owns final decisions. It also pre-empts a likely misreading: passing agent_profile never selects a session profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core behavior in the first sentence, then usage guidance, then scope constraints. Slightly dense and broken across lines, but every sentence carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, yet the description still characterizes the checklist's contents. Combined with the statelessness and session-selection caveats, an agent has everything needed to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum description already maps each profile to its concern, so the baseline is 3. The description adds real semantic value by clarifying what agent_profile does not do (it never selects a session profile), which is not derivable from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return a fixed checklist for one agent_profile') and enumerates what the checklist contains (review focus, required output sections, shared evidence constraints, escalation signals). It is clearly distinguishable from the session-lifecycle siblings (start/advance/finish/get_qa_orchestration).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes usage to 'one scoped concern' and names the alternative path: request each review_profiles member separately in session order for a bundle, and use start_qa_orchestration to track the multi-concern review. Both when-to-use and the sibling alternative are spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_qa_orchestrationA

Create a content-free QA session for a tracked review; call once before triage.

Every call creates a separate session. Use prepare_qa_orchestration for a stateless profile checklist. task_type narrows the bundle shortlist; the host still selects the review route.

  • Sessions expire after the configured TTL (1800 seconds by default).

  • Capacity is 100 sessions by default and can be configured.

  • Expired sessions are purged and retained terminal sessions may be evicted. If capacity can only be freed by removing an active session, the call returns a session limit error.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_typeYesBroad request category: ordinary_review for implementation review, widget_review for widget changes, epic_analysis for broad exploration, requirements_analysis for requirement review, qa_planning for coverage planning, autotest_implementation for automation work, or other. This affects recommended_bundles only; the host chooses the route from changed files and evidence; it is not a risk level.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover safety/idempotency flags; the description adds substantial behavioral context beyond them: TTL expiry (1800s default), 100-session capacity, purge/eviction behavior, and the specific 'session limit' error when only an active session could free capacity. This is exactly the operational detail annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence, then a routing sentence, then tight bulleted constraints. Every sentence carries distinct information with no filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema covers return values, so the description is not obliged to describe them. Combined with annotations and the schema, the description supplies the lifecycle (TTL, capacity, eviction, error) an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum is fully self-documented, so baseline 3 applies. The description still adds real meaning by clarifying what task_type does and does not do ('narrows the bundle shortlist', 'host still selects the review route', 'not a risk level'), correcting a plausible misreading.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create) and resource (content-free QA session) with the timing constraint (call once before triage) front-loaded. It explicitly contrasts with prepare_qa_orchestration, so the agent can separate it from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when ('call once before triage') and when-to-use-the-alternative ('Use prepare_qa_orchestration for a stateless profile checklist'). It also clarifies that every call creates a separate session, preempting accidental reuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.12
    • Addedprepare_qa_orchestration
    • Removedprepare_review_route
  2. 1 tool updatev0.1.6
    • Changedfinish_qa_orchestration1 field changed
      • changedInput schema / properties / outcome / description
        Previous value: -"Host-selected final QA outcome: completed after synthesis, or the same partial or blocked outcome already used for an early stop. Conflicting final outcomes are rejected."New value: +"Host-selected final session status: completed, partial, or blocked."
  3. 5 tool updatesv0.1.1
    • Changedadvance_qa_orchestration8 fields changed
      • addedInput schema / properties / completed_profile / description
        Added value: +"After each successful primary-review step, submit exactly the current_profile returned by the previous advance_qa_orchestration call. Omit for other steps or an early stop."
      • addedInput schema / properties / completed_step / description
        Added value: +"Active step returned in current_step and completed by this call: triage, primary_review, deep_review, or synthesis. When current_step is awaiting_host_outcome, call finish_qa_orchestration instead."
      • changedInput schema / properties / completed_step / enum
        Previous value: -[
        -  "triage",
        -  "primary_review",
        -  "deep_review",
        -  "synthesis",
        -  "awaiting_host_outcome"
        -]New value: +[
        +  "triage",
        +  "primary_review",
        +  "deep_review",
        +  "synthesis"
        +]
      • changedInput schema / properties / risk_signals / description
        Previous value: -"Required with the final completed primary-review profile. Send an empty object when no signals apply."New value: +"Required with the final completed primary-review profile. Send an empty object when no evidence-based fixed escalation signals apply; do not send on earlier profiles or on the synthesis transition."
      • addedInput schema / properties / run_id / description
        Added value: +"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
      • addedInput schema / properties / selected_bundle / description
        Added value: +"For completed triage, choose one allowed fixed bundle for a broad or cross-concern review. It is mutually exclusive with selected_profile; recommended_bundles is only a shortlist."
      • addedInput schema / properties / selected_profile / description
        Added value: +"For completed triage, choose one allowed profile for a narrow, low-risk review. It is mutually exclusive with selected_bundle."
      • addedInput schema / properties / status / description
        Added value: +"Use completed after the active step succeeds. Use partial or blocked to stop early; then finalize with the same outcome using finish_qa_orchestration."
    • Changedfinish_qa_orchestration2 fields changed
      • addedInput schema / properties / outcome / description
        Added value: +"Host-selected final QA outcome: completed after synthesis, or the same partial or blocked outcome already used for an early stop. Conflicting final outcomes are rejected."
      • addedInput schema / properties / run_id / description
        Added value: +"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
    • Changedget_qa_orchestration1 field changed
      • addedInput schema / properties / run_id / description
        Added value: +"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
    • Changedprepare_review_route1 field changed
      • addedInput schema / properties / agent_profile / description
        Added value: +"Choose one fixed specialist profile: code_explorer maps execution and evidence; code_reviewer reviews implementation; pr_test_analyzer reviews coverage; security_reviewer and silent_failure_hunter target security and silent failures; ruby_reviewer, python_reviewer, typescript_reviewer, react_reviewer, and mobile_reviewer target their named stacks. Choose from changed files and confirmed stack."
    • Changedstart_qa_orchestration1 field changed
      • addedInput schema / properties / task_type / description
        Added value: +"Broad request category: ordinary_review for implementation review, widget_review for widget changes, epic_analysis for broad exploration, requirements_analysis for requirement review, qa_planning for coverage planning, autotest_implementation for automation work, or other. This affects recommended_bundles only; the host chooses the route from changed files and evidence; it is not a risk level."
  4. 5 tool updatesv0.1.0
    • First observedadvance_qa_orchestration
    • First observedfinish_qa_orchestration
    • First observedget_qa_orchestration
    • First observedprepare_review_route
    • First observedstart_qa_orchestration

TDQS

A4.7/5.0

Scored across 5 tools

Disambiguation4/5

The five tools map to distinct lifecycle stages (prepare=stateless lookup, start=create session, advance=progress step, finish=terminal outcome, get=read state). The main risk is confusing prepare_qa_orchestration with start_qa_orchestration, but the descriptions explicitly disambiguate by stressing prepare is stateless and never creates a session. advance vs finish is also clarified by tying finish to the awaiting_host_outcome terminal state.

Naming Consistency5/5

All five tools follow a strict verb_qa_orchestration pattern: start_, prepare_, advance_, finish_, get_. The convention is uniform across the entire set with no deviations or mixed casing.

Tool Count5/5

Five tools is a tight, well-scoped set that exactly covers a session lifecycle plus a stateless helper. Each tool earns its place with no redundancy or padding.

Completeness4/5

The surface covers the full lifecycle (create, stateless lookup, progress, read, terminal finalize) and handles early stops. Minor gap: there is no explicit list/abort/cancel operation, though early stop is folded into advance + finish, so agents can still work around it.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers