Skip to main content
Glama

QA Orchestrator

QA Orchestrator MCP server – quality and maintenance score on Glama

QA Orchestrator is a small deterministic FastMCP service for host-owned QA reviews. It keeps orchestration state bounded; evidence, source code, logs, prompts, model responses, and final decisions remain with the primary host agent.

Quick start

Requires Python 3.12+, uv, Git, and a local MCP client with STDIO support.

git clone https://github.com/Rbkmen/qa-orchestrator.git
cd qa-orchestrator
uv sync --locked
uv run qa-orch setup
uv run qa-orchestrator-doctor

Then register the server and add the host instructions. The MCP client starts the server. To inspect the models loaded by that process, call get_qa_orchestration_model_policy with {} in the client. See the installation guide for Windows, Claude Code, installation without a checkout, updates, and troubleshooting.

Related MCP server: Support Ticket Triage MCP

How it works

  1. The primary host obtains authoritative evidence from the required systems and classifies the QA task.

  2. The host calls start_qa_orchestration. The orchestrator creates a content-free session and returns the first step and its configured policy. Run qa-orch setup to choose OpenAI or Anthropic and configure models and reasoning for each stage.

  3. The host runs each stage in its configured model environment and sends the orchestrator only a structured signal after each stage:

    • triage — select one fixed review bundle or one compatibility profile;

    • primary review — review every selected profile in the fixed order;

    • optional deep review — perform one read-only analysis when fixed risk signals match;

    • final synthesis — consolidate the results. Each stage uses the model and reasoning configured for it in the returned model_policy. The returned policy selects the provider, model, and reasoning for each stage. Execution speed and latency preferences remain controlled by the user's host/provider settings; the orchestrator does not set or override them.

  4. The host validates findings, runtime evidence, and limitations. For an orchestrated task, it calls finish_qa_orchestration with the same run_id and its final outcome.

The orchestrator does not call models, choose severity or release readiness, or perform external writes. Its client-rule templates provide portable host-agent instructions for evidence quality and finding presentation; registering the MCP server alone does not load those instructions into the host.

Choose a review path

Use one compatibility profile by default for a routine, narrowly scoped, low-risk change with one main concern. Use a fixed bundle for broad, cross-concern, or high-risk changes. For example, a small Ruby guard change can use ruby_reviewer; a change spanning a Rails endpoint, a background job, and their tests fits ruby_backend.

recommended_bundles is a task-type shortlist, not a risk score or a required selection. The host inspects the diff and chooses the initial profile or bundle. The host sends structured risk signals after primary review; the orchestrator applies its fixed rules to decide whether optional deep review follows.

Example: the same review with and without the orchestrator

Ordinary review guided by AGENTS.md

Review with QA Orchestrator

Small Ruby guard change

The host chooses a reviewer prompt and tracks the review in the conversation.

The host selects ruby_reviewer; the MCP validates the transition and returns the bounded session state.

Broad Ruby change across an endpoint, job, and tests

The host coordinates review steps from its instructions.

The host selects ruby_backend; the fixed reviewer order and required transitions are explicit.

Deeper review

The host decides from its own instructions and evidence.

The host sends structured risk signals; the orchestrator applies fixed escalation rules.

Ownership

The host gathers evidence and makes the final decision.

The host still owns evidence, model calls, findings, and the final decision; the orchestrator receives no source or raw evidence.

The orchestrator adds a validated workflow contract and bounded progress state. It does not replace host instructions or perform the review itself.

QA Orchestrator at a glance

Detailed host-led QA Orchestrator workflow, including review stages, optional escalation, and ownership boundaries.

The diagram shows default memory-only operation. Optional SQLite recovery is described in the MCP interface.

Bundles and profile names

The triage stage selects one fixed bundle or one compatibility profile. start_qa_orchestration returns recommended_bundles as a task-type-based shortlist; it does not restrict allowed_bundles. Choose the route from the changed files and confirmed project stack. The primary-review stage executes bundle profiles sequentially. After every role, the host sends completed_profile, and the orchestrator returns current_profile and completed_profiles. Use the updated session returned by advance_qa_orchestration for the next action and model policy; call get_qa_orchestration only when resuming or recovering a session. Deep review or synthesis is available only after the final role. Transition identifiers are model-neutral; choose the model from the returned model_policy, never from the step name.

Bundle

Profile order

ordinary_mr

code_explorer → code_reviewer → pr_test_analyzer

widget

code_explorer → react_reviewer → typescript_reviewer → pr_test_analyzer

widget_js

code_explorer → code_reviewer → react_reviewer → pr_test_analyzer

ruby_backend

code_explorer → ruby_reviewer → pr_test_analyzer

python_backend

code_explorer → python_reviewer → pr_test_analyzer

mobile

code_explorer → mobile_reviewer → pr_test_analyzer

security

code_explorer → security_reviewer → silent_failure_hunter

autotest

code_reviewer → pr_test_analyzer → typescript_reviewer

requirements

code_explorer → code_reviewer

autotest and widget include a TypeScript review role; select them when TypeScript review is relevant. Use ordinary_mr for broad non-TypeScript automation, widget_js for broad JavaScript React changes, ruby_backend for broad Ruby backend changes, python_backend for broad Python/MCP changes, and mobile for broad React Native or native iOS/Android changes. Choose from the changed files and confirmed project manifests, not the repository name alone; monorepos can contain several stacks.

Technical profiles and display names:

Profile

Host-facing name

code_explorer

Faraday — Evidence Investigator

code_reviewer

Code Reviewer

pr_test_analyzer

Test Analyzer

security_reviewer

Security Reviewer

silent_failure_hunter

Silent Failure Hunter

typescript_reviewer

TypeScript Reviewer

react_reviewer

React Reviewer

ruby_reviewer

Ruby Reviewer

python_reviewer

Python Reviewer

mobile_reviewer

Mobile Reviewer

Faraday is only the internal display name of the code_explorer profile. No external agent, service, package, or model is connected under that name.

Ordinary MR flow with optional escalation:

Triage → Ordinary MR Review
Primary review → Faraday — Evidence Investigator → Code Reviewer → Test Analyzer
  ├─ no escalation ───────────────────────────────→ Final synthesis
  └─ fixed risk signals match → Deep review → Final synthesis
Host → Final QA outcome

With the final primary-review role, the host must send the structured boolean risk_signals object in the same advance_qa_orchestration call as the final completed_profile (send {} when no signals apply). Deep review is derived only from the fixed signal rules. Do not send risk_signals on the later synthesis transition. The returned session includes matched rules and fixed reason codes as deep_assessment; raw evidence never enters the orchestrator.

Deep-review rules:

  1. high_risk_domain + evidence_uncertain;

  2. any two of cross_system_scope, multiple_plausible_causes, non_reproducible, and high_blast_radius;

  3. evidence_conflict together with high_risk_domain, cross_system_scope, or high_blast_radius.

For low-risk, narrow reviews, the triage stage may select one compatibility profile instead of a bundle: code_reviewer for a small behavior change, pr_test_analyzer for a test-only change, typescript_reviewer for a TypeScript-only change, react_reviewer for a React-only change, ruby_reviewer for a Ruby-only change, python_reviewer for a Python/MCP-only change, or mobile_reviewer for a React Native/native-platform-only change. Broad or cross-concern reviews continue to use a fixed bundle.

Keep one compact per-task Evidence Packet with stable evidence references (E1, E2, ...) and bounded finding candidates (F-01, F-02, ...). Do not repeat the full diff or raw logs in every model stage.

MCP interface

The service publishes exactly nine tools:

Tool

Purpose

prepare_qa_orchestration(agent_profile)

Return the fixed checklist for one profile without creating a session

get_qa_orchestration_catalog()

Discover all profiles, bundle purposes and ordered routes, and task-type shortlists without creating a session

get_qa_orchestration_model_policy()

Read the server-wide provider, models, and reasoning for all stages of new sessions; no run_id

start_qa_orchestration(task_type)

Create a session and return the task-based bundle shortlist

advance_qa_orchestration(...)

Record one active review step's completion or an early stop; return the next action

get_qa_orchestration(run_id)

Read one existing session's step, status, next action, and current-stage model policy

list_qa_orchestrations()

Recover lost run_id values from brief metadata for non-expired sessions retained by this server

finish_qa_orchestration(run_id, outcome)

Finalize the host-owned session outcome after synthesis or a recorded early stop

delete_qa_orchestration(run_id)

Immediately discard one session's in-memory state and its configured recovery record

Before starting a session or choosing its triage route, call get_qa_orchestration_catalog with {}. It returns all available profiles with their display names and focus, all bundles with usage guidance and ordered profile IDs, and recommended_bundles_by_task_type. Recommendations are shortlists, not restrictions; choose from changed files and confirmed stack. For a narrow concern, choose one profile and get its detailed checklist through prepare_qa_orchestration. For broad work, select a bundle and preserve its profile order. The catalog is fixed, independent of session/model-policy state, and does not create sessions, change TTL, or write storage.

Bundle orchestration flow:

Triage → Primary review[1] → ... → Primary review[N]
                                      ↘ optional Deep review ↗
                                           Final synthesis → Host outcome

Sessions are kept in process memory by default. Successful state changes refresh the 1,800-second default TTL; reads do not. The maximum is 100 active sessions, and the shared cache is bounded, so older terminal sessions may be evicted when capacity is needed. Repeating the final call is idempotent while its session is retained. After a restart, the host starts a new session unless optional recovery storage is enabled. read_only=true and host_owns_decisions=true are part of every state.

Set QA_ORCHESTRATOR_SESSION_STORE_PATH to opt into a local SQLite file that restores unfinished orchestration state after a restart. It stores only the current structured session needed for recovery; finalization removes that row. It never stores evidence, prompts, source, logs, model responses, finalized outcomes, history, or statistics. Use one server process per store file. The default remains memory-only.

If a run_id is lost, call list_qa_orchestrations with {}. Its sessions entries include the ID, task type, status, stage, selected route, current profile, and expiry; use these to identify the intended session, then call get_qa_orchestration with its ID. Confirm the intended session if several entries match. The list includes retained terminal sessions and is sorted by expiry descending, with run_id as tie-breaker. It does not change state, renew TTL, or write storage. An empty list means no non-expired sessions are retained by this server; expired or evicted sessions cannot be recovered. After a restart, only unfinished sessions restored from configured recovery storage are available.

After synthesis, the session waits for the host's final outcome. Call finish_qa_orchestration with the session run_id and completed, partial, or blocked. For an early stop, first pass partial or blocked to advance_qa_orchestration, then finish the session with the same outcome. Repeating the same finalization is idempotent; a conflicting outcome is rejected. Tasks that do not use orchestration need no finalization call. The service does not store or report task statistics.

To intentionally discard a session, call delete_qa_orchestration with its run_id. It removes only that session from memory and optional SQLite recovery storage, at any lifecycle stage, and frees its capacity immediately. It returns { "run_id": "qar-...", "deleted": true }; deleted confirms absence, so an already missing or expired ID also succeeds. The ID can no longer be read, advanced, finalized, or restored after restart. Storage failure leaves memory state intact and can be retried. Deletion records no QA outcome and does not stop host tasks or model executions. Use normal early-stop/finalization when the outcome should remain available; delete only when the host intends to discard that session.

Review profiles

prepare_qa_orchestration returns the focus, required sections, constraints, escalation signals, and display name for one profile. For a bundle, the host calls the route for every profile in the returned fixed order and passes its technical identifier in completed_profile after each call.

required_sections is profile-specific: the evidence investigator returns Scope, Evidence Map, and Unverified; test analysis returns Scope, Coverage Gaps, and Unverified; implementation and specialist reviews return Scope, Finding Candidates, Coverage Gaps, and Unverified.

Use the local profile evaluation pack to smoke-check role boundaries and bundle selection without collecting task statistics.

Responsibility boundary

The primary host is responsible for:

  • obtaining and validating evidence;

  • calling the issue tracker, code host, test-management system, observability and logging systems, documentation and chat systems, code index, and the file system;

  • running triage, primary-review, synthesis, and any optional deep-review stage under the policy;

  • confirmed findings, severity, release/readiness judgment, and the final QA response;

  • file changes and all external writes.

QA Orchestrator is responsible only for fixed routing, state transitions, read-only constraints, and bounded session finalization. advance_qa_orchestration must not receive an Evidence Packet, prompt, model output, source text, logs, paths, or an arbitrary reason.

The read_only=true flag describes the review boundary. MCP annotations mark start, advance, and finish as state-changing tools because they update local orchestration state. Profile lookup, model-policy inspection, and state lookup are annotated as read-only.

Requirements

  • Python 3.12+;

  • uv;

  • an MCP client that supports STDIO;

  • Windows, macOS, or Linux.

On Windows, uv sync installs the native MCP entry point at .venv\Scripts\qa-orchestrator-mcp.exe. macOS and Linux use the POSIX source launcher.

Installation

For the complete setup—including Codex and Claude Code registration, host instructions, verification, and troubleshooting—see the installation guide.

Start with the quick start above to install from a checkout.

On macOS and Linux, the source launcher uses the project's .venv, the active VIRTUAL_ENV, or an installed qa-orchestrator-mcp from PATH. On Windows, use the installed .venv\Scripts\qa-orchestrator-mcp.exe entry point. No separate background process is required.

For clients that support uvx, a checkout is optional:

Replace <commit-sha> with the full commit hash and use the same hash in the setup and server commands.

uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' qa-orch setup

This saves the model policy in your user configuration. Register the server command with your MCP client so the client starts it when needed; do not run the server command directly in a terminal. For Codex without a checkout:

codex mcp add qa-orchestrator -- uvx \
  --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' \
  qa-orchestrator-mcp

Use an immutable commit SHA instead of the default branch for reproducible team configuration. See the installation guide for the remaining no-checkout commands and update instructions.

The setup wizard first lets you choose Russian or English for that run, then stores only the provider label, model IDs, and the selected provider-specific reasoning/effort values in the local model-policy.json; the language is not saved, and the wizard never asks for or stores API keys. Choose a provider and model that your host client can use; the wizard records the policy but does not configure provider access. From a checkout, use uv run qa-orch config show and uv run qa-orch reload. Without a checkout, prefix those commands with uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>'. Colors in the wizard distinguish providers, model IDs, and reasoning values; set NO_COLOR=1 to disable them or FORCE_COLOR=1 to force them.

The wizard uses a local recommendation/capability catalog and does not contact provider APIs. For custom model IDs, verify that the host account can access the model and supports the selected reasoning/effort value.

Inspect and apply model settings

Check

What it proves

uv run qa-orch config show

The saved policy on disk, or built-in defaults if no file exists

uv run qa-orch reload

The CLI can read and validate that file; the connected server still needs a restart

MCP get_qa_orchestration_model_policy({})

The policy currently loaded by the connected server for new sessions

Host execution details

Which model actually performed a review stage

After changing settings, restart the client's server connection and call the MCP policy tool again. The orchestrator provides policy metadata; it cannot switch the host's model or verify that the host executed that model. If the host cannot use a requested model, report that limitation in the review.

Without a saved policy, defaults are gpt-6-luna / max for triage and gpt-6-sol for primary review / medium, deep review / high, and synthesis / medium. Setup preserves an existing selection when you accept its defaults.

Example for Codex with a local checkout:

codex mcp add qa-orchestrator -- "$(pwd)/scripts/qa-orchestrator"

On Windows, run this from the repository root in PowerShell:

codex mcp add qa-orchestrator -- "$PWD\.venv\Scripts\qa-orchestrator-mcp.exe"

The two supported host integrations are described in the client guides.

Configuration

Variable

Default

QA_ORCHESTRATOR_DATA_DIR

~/.qa-orchestrator

QA_ORCHESTRATOR_ORCHESTRATION_TTL_SECONDS

1800

QA_ORCHESTRATOR_ORCHESTRATION_MAX_SESSIONS

100

QA_ORCHESTRATOR_MODEL_POLICY_PATH

~/.qa-orchestrator/model-policy.json

QA_ORCHESTRATOR_SESSION_STORE_PATH

unset (disabled)

The default policy path follows QA_ORCHESTRATOR_DATA_DIR. Configure the same policy path for setup and for the MCP client; use absolute paths for portable client configuration. Limits must be positive integers. See the configuration example for client environment overrides.

Finalize an orchestration

For completion and early-stop instructions, see the MCP interface section.

Clients and rules

Development

See CONTRIBUTING.md. Any change to the public MCP contract must include an exact tool-surface test and a check that evidence, decisions, and external writes remain with the primary host.

Available Tools

9 tools
advance_qa_orchestrationRecord QA step resultA

Record one active QA step's result; return updated state and next_action instructions for the host.

For finalization, use finish_qa_orchestration at awaiting_host_outcome; after early stops, match the recorded outcome.

  • Triage: select exactly one bundle or profile.

  • Primary review: follow profile order; the final profile requires risk_signals ({} if none).

  • Early stop: omit route selectors, completed_profile, and risk_signals.

  • Success saves state and refreshes expires_at.

  • Invalid/stale requests and unknown/expired run_ids fail without advancing.

  • Non-idempotent: after a lost response, read get_qa_orchestration(run_id); prior-step replay is rejected.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesSession identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it.
statusYesUse completed after the active step succeeds. Use partial or blocked to stop early; then finalize with the same outcome using finish_qa_orchestration.
risk_signalsNoRequired with the final completed primary-review profile. Send an empty object when no evidence-based fixed escalation signals apply; do not send on earlier profiles or on the synthesis transition.
completed_stepYesActive step returned in current_step and completed by this call: triage, primary_review, deep_review, or synthesis. When current_step is awaiting_host_outcome, call finish_qa_orchestration instead.
selected_bundleNoFor completed triage, choose one allowed fixed bundle for a broad or cross-concern review. It is mutually exclusive with selected_profile; recommended_bundles is only a shortlist.
selected_profileNoFor completed triage, choose one allowed profile for a narrow, low-risk review. It is mutually exclusive with selected_bundle.
completed_profileNoAfter each successful primary-review step, submit exactly the current_profile returned by the previous advance_qa_orchestration call. Omit for other steps or an early stop.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the mutation/non-idempotent profile, but the description goes further: state is saved and expires_at refreshed on success, invalid/stale requests and unknown/expired run_ids fail without advancing, and prior-step replay is rejected. This is exactly the failure-mode and retry context the annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and return value in one sentence, then tightly scoped bullets for triage, primary review, early stop, success, failure, and retry. Every line encodes a distinct rule; no filler or repetition of schema text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, non-idempotent orchestration step with an output schema, the description covers routing, per-step parameter requirements, failure semantics, and recovery, and it correctly does not re-explain return values. An agent has everything needed to call it correctly on the first attempt.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds cross-parameter conditional logic the schema does not express: risk_signals is required only with the final completed primary-review profile ({} when none), completed_profile must echo current_profile, and route selectors must be omitted on early stop. It does not restate the enum values themselves, which is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Record one active QA step's result") plus the return contract (updated state and next_action). It explicitly distinguishes itself from the two nearest siblings, finish_qa_orchestration and get_qa_orchestration, so an agent can route without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when/when-not rules and alternatives: use finish_qa_orchestration at awaiting_host_outcome, match the recorded outcome after early stops, do not call on a lost response but read get_qa_orchestration(run_id) instead. Per-step selection rules (triage picks exactly one bundle or profile; early stop omits selectors) leave nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_qa_orchestrationDiscard QA session stateA
DestructiveIdempotent

Delete one QA session from memory and configured recovery storage. To retain a final QA outcome, use finish_qa_orchestration.

run_id selects one session at any stage; no prior advance or finish required. Recover a lost ID via list_qa_orchestrations(), matching task type, route, and stage; confirm the target if several entries match.

  • Idempotent: deleted=true confirms absence, including unknown or expired IDs.

  • Irreversible: get/advance/finish cannot use the ID, even after restart.

  • Storage failure: memory stays intact; resolve the error and retry.

Host tasks and model executions are unaffected.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesSession identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
deletedNoThe run_id is absent from this server and its recovery store, including if already absent.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive, idempotent, and non-read-only, but the description adds material context beyond them: deleted=true confirms absence even for unknown/expired IDs, irreversibility that survives restart, storage-failure behavior where memory stays intact, and the explicit note that host tasks and model executions are unaffected. These are exactly the operational details annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded imperative sentence, then a bulleted block where each line carries a distinct fact (mutation behavior, idempotency, irreversibility, failure recovery, unaffected systems). No sentence restates the name or the annotations verbatim, and there is no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return fields, and it correctly stops there. For a single-parameter destructive mutation with full annotation coverage, everything an agent needs to call it safely and recover from edge cases is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema itself documents the qar- pattern, so the baseline is 3. The description nonetheless adds dispatch-relevant meaning: run_id can select a session at any stage with no prior advance or finish, and it points at list_qa_orchestrations as the recovery path for an unknown ID.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Delete one QA session from memory and configured recovery storage') and immediately distinguishes itself from the sibling that shares its domain by naming finish_qa_orchestration as the retain-outcome alternative. An agent can differentiate it from advance/get/list/prepare with no schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when to use this tool ('to retain a final QA outcome, use finish_qa_orchestration' implies deletion is for discarding), relaxes a constraint ('no prior advance or finish required'), and routes the agent to list_qa_orchestrations to recover a lost ID with instructions for disambiguating matches. Explicit alternatives and conditions are provided rather than implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

finish_qa_orchestrationRecord final QA outcomeA
Idempotent

Finalize the whole QA run's outcome; active step results belong to advance_qa_orchestration.

After synthesis, use the retained run_id at awaiting_host_outcome to choose completed, partial, or blocked. After an early stop recorded by advance_qa_orchestration, outcome must match its partial/blocked status.

  • Same outcome: returns the retained terminal session.

  • Different outcome: conflicting final outcome; status unchanged.

  • Premature call: outcome is not ready. Finalized sessions expire at configured TTL (1800s default), or are removed by eviction or restart. Removed run_ids cannot be deduplicated.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesSession identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it.
outcomeYesHost-selected final session status: completed, partial, or blocked.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=false). It discloses idempotent semantics concretely (same outcome returns the retained terminal session; a different outcome raises 'conflicting final outcome' with status unchanged), the TTL expiry (1800s default), eviction/restart removal, and that removed run_ids cannot be deduplicated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then bulleted behavioral cases (same outcome / different outcome / premature call) and a closing lifecycle sentence. Dense but every sentence carries distinct information with no repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description instead covers the error paths, state preconditions, and session lifecycle that an agent must know to call this correctly. Nothing material is missing for a two-parameter finalization tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: run_id must be the retained session id from start_qa_orchestration and the outcome choice is constrained by upstream early-stop state. The enum values themselves are already documented in the schema, so it does not fully re-earn a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Finalize the whole QA run's outcome') and immediately scopes it against the sibling that handles the other half of the work ('active step results belong to advance_qa_orchestration'). An agent can distinguish it from advance_qa_orchestration, start_qa_orchestration and the get/list siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit preconditions and timing: call after synthesis, use the retained run_id while the session is at awaiting_host_outcome, and ensure the outcome matches a partial/blocked early stop recorded by advance_qa_orchestration. It also warns that a premature call returns 'outcome is not ready', so when-not-to-call is covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_qa_orchestrationInspect one QA sessionA
Read-onlyIdempotent

Read one retained run_id's content-free state: status, current step, next action, and current-stage model policy.

Use run_id on the server retaining that session to resume or reconcile a lost advance_qa_orchestration response. Reads do not extend expires_at. Errors unknown run_id or expired session require a new session via start_qa_orchestration. If run_id is lost, use list_qa_orchestrations() to find retained sessions first.

For all-stage server policy, use get_qa_orchestration_model_policy(); for review progress or final outcome, use advance_qa_orchestration or finish_qa_orchestration, respectively. Local stdio access requires no additional credentials.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesSession identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag readOnly/idempotent/closed-world, and the description adds genuinely new behavior: reads do not extend expires_at, the id is server-scoped, and local stdio access needs no extra credentials. These are exactly the operational facts an agent cannot infer from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what is read, then scoping, then error/recovery routing, then disambiguation from siblings. Every sentence carries distinct information; none is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no further explanation, and the description still summarizes the returned state. For a single-parameter read tool, nothing an agent needs to call it correctly or recover from failure is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the pattern plus reuse warning are already in the schema, so baseline is 3. The description adds real meaning by specifying the id must be used on the server retaining that session and that it can be recovered from list_qa_orchestrations, which the schema does not state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (one run_id's retained state) and enumerates exactly what is returned: status, current step, next action, model policy. It is immediately distinguishable from sibling readers like get_qa_orchestration_model_policy and list_qa_orchestrations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly covers when to use it (resume or reconcile a lost advance_qa_orchestration response), what the error cases mean (`unknown run_id`, `expired session`) and the recovery path (start_qa_orchestration), and what to do if the id is lost (list_qa_orchestrations). It also names the alternatives for other needs (model policy, review progress, final outcome).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_qa_orchestration_catalogDiscover QA profiles and bundlesA
Read-onlyIdempotent

Read the static profile/bundle catalog: available specialists, review scopes, and ordered routes.

Use before creating a session or choosing its triage route; no session or arguments required. Task-type recommendations are shortlists, not restrictions. Choose a bundle for broad work or a single profile for a narrow concern using changed files and confirmed stack.

  • Detailed checklist: prepare_qa_orchestration(agent_profile).

  • Create a session: start_qa_orchestration; set its route: advance_qa_orchestration.

  • Session IDs: list_qa_orchestrations().

Available routes come from the installed server's bundled definitions; active sessions and configured models do not change this catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
bundlesYes
profilesYes
read_onlyNo
host_owns_decisionsNo
recommended_bundles_by_task_typeYesTask-based shortlists only; they do not restrict the available bundles.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false. The description adds useful context beyond them: the catalog is static and comes from the installed server's bundled definitions, active sessions and configured models do not change it, and task-type recommendations are shortlists rather than restrictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose, then moves into usage conditions, cross-references, and a final caveat. The bullet list for related tools is efficient and nothing reads as filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only discovery tool with an output schema, the description provides everything an agent needs: what it returns, when to call it, what preconditions are unnecessary, and how it relates to sibling tools. Return-value details are appropriately left to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the rubric baseline is 4. The description adds a small amount of useful clarity by stating that no session or arguments are required, though there is no parameter syntax or structure to explain further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: reading the static profile/bundle catalog, with its contents enumerated (specialists, review scopes, ordered routes). It is clearly distinguishable from siblings like prepare_qa_orchestration and start_qa_orchestration because it describes discovery, not session mutation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use this before creating a session or choosing its triage route, and notes that no session or arguments are required. It also provides next-step alternatives for detailed checklists, session creation, route setting, and listing session IDs, and explains how to choose a bundle versus a single profile.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_qa_orchestration_model_policyInspect server model settingsA
Read-onlyIdempotent

Read server-wide model settings for every stage: provider, model IDs, and reasoning.

Use to inspect the policy loaded for new sessions; takes no arguments.

  • Existing session state and current-stage policy: get_qa_orchestration(run_id).

  • Create a session with these settings: start_qa_orchestration.

Reads the loaded in-memory snapshot, not the current policy file; restart the server connection after file changes. No session is created and no provider is contacted; the result does not verify which model the host executed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
providerYes
deep_modelYes
triage_modelYes
primary_modelYes
deep_reasoningNo
synthesis_modelYes
triage_reasoningNo
primary_reasoningNo
synthesis_reasoningNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and closed-world behavior, but the description adds substantial non-obvious context: it reads an in-memory snapshot rather than the current policy file, requires a server restart after file changes, creates no session, contacts no provider, and does not verify which model the host executed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and scope, then concisely groups usage routing into bullets and behavioral caveats into tight sentences. Every sentence adds distinct value with no repetition or wasted space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return values need not be described. The description covers what is read, when to use it, alternatives, and important behavioral caveats, making it complete for a zero-argument inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description confirms 'takes no arguments,' which is consistent with the empty schema and leaves no parameter ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: reads server-wide model settings for every stage, listing provider, model IDs, and reasoning. It also distinguishes itself from siblings by explicitly naming get_qa_orchestration(run_id) and start_qa_orchestration as alternative tools for different needs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it to inspect the policy loaded for new sessions and notes it takes no arguments. It routes the agent clearly: existing session/current-stage policy uses get_qa_orchestration, and creating a session with these settings uses start_qa_orchestration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_qa_orchestrationsList retained QA sessionsA
Read-onlyIdempotent

List non-expired QA session identifiers and brief metadata retained by this server.

Recover a lost run_id by matching task type, route, and stage; confirm if several match. Call get_qa_orchestration(run_id) for full state and next action, or get_qa_orchestration_model_policy() for server policy. Takes no arguments.

  • Includes retained terminal sessions; expired or evicted sessions are unrecoverable.

  • Sorts by expires_at descending, then run_id descending; empty means none are retained.

  • After restart, only unfinished sessions restored from configured local storage appear.

  • Reading does not extend session TTL.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
sessionsYes
read_onlyNo
host_owns_decisionsNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/openWorldHint=false, so the safety profile is covered; the description goes further with non-obvious behavior the schema cannot express: reading does not extend TTL, expired/evicted sessions are unrecoverable, sorting is expires_at desc then run_id desc, empty result means nothing retained, and post-restart only unfinished local-storage sessions reappear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence states purpose, the second sentence front-loads the primary use case and alternatives, and four tight bullets carry only non-redundant behavioral facts (TTL, sort, restart, empty semantics). No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists (so return-field documentation is not required), the description still supplies everything an agent needs that the structured fields do not: empty-result interpretation, sort order, TTL non-extension, and restart scoping. Nothing material to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and schema description coverage is 100%, so the baseline is 4. The description explicitly confirms "Takes no arguments," which is the only parameter-level fact available and is already fully specified by the schema, so it cannot exceed the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List non-expired QA session identifiers and brief metadata") and immediately distinguishes itself from get_qa_orchestration, which returns "full state" rather than a list. An agent can tell from the description alone which of the siblings to call for a lookup versus a state read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit use case ("Recover a lost run_id by matching task type, route, and stage; confirm if several match") and routes to the correct alternatives by name for other needs (get_qa_orchestration for full state, get_qa_orchestration_model_policy for server policy). It also states a when-not: expired or evicted sessions are unrecoverable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_qa_orchestrationRead specialist review checklistA
Read-onlyIdempotent

Return one specialist's fixed ReviewRoute checklist: profile, display_name, focus, required_sections, constraints, and escalation_signals.

Use before reviewing one scoped concern; no session is required.

  • Standalone review: use get_qa_orchestration_catalog() to match profile focus to changed files and confirmed stack.

  • Session primary_review: pass the latest current_profile as agent_profile; preserve review_profiles order for bundles.

This local lookup reads bundled definitions, without executing reviews or calling models. Use advance_qa_orchestration to apply a session route. Results depend only on agent_profile. Invalid values fail input validation; retry with a listed enum value.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_profileYesChoose one fixed specialist profile: code_explorer maps execution and evidence; code_reviewer reviews implementation; pr_test_analyzer reviews coverage; security_reviewer and silent_failure_hunter target security and silent failures; ruby_reviewer, python_reviewer, typescript_reviewer, react_reviewer, and mobile_reviewer target their named stacks. Choose from changed files and confirmed stack.

Output Schema

ParametersJSON Schema
NameRequiredDescription
focusYesThe profile's review boundary and primary responsibility.
profileYes
read_onlyNo
constraintsYesShared limits for this read-only review profile.
display_nameYes
required_sectionsYesProfile-specific output sections; use these instead of a generic template.
escalation_signalsNoConditions to assess against evidence. A listed condition is not itself a deep-review trigger; set risk_signals only when the corresponding evidence-based condition applies.
host_owns_decisionsNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/openWorldHint/idempotentHint annotations, it discloses that this is a local lookup over bundled definitions that executes no reviews and calls no models, that results are a pure function of agent_profile, and that invalid values fail input validation and should be retried with a listed enum value. This error-handling and no-side-effect disclosure is genuinely non-obvious context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by tightly scoped bullets for the two usage paths and a short closing note on local behavior and validation. Every sentence carries information; nothing is redundant with the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one required enum parameter, a 100%-covered schema, and an output schema present, the remaining burden is routing and behavior – both covered (catalog vs. session path, advance_qa_orchestration alternative, determinism, validation failure). Nothing an agent needs in order to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum description already defines all ten profiles, so the schema carries the baseline. The description still adds selection guidance not in the schema: results depend only on agent_profile, choose from changed files and confirmed stack, and preserve review_profiles order when handling bundles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb and resource ('Return one specialist's fixed ReviewRoute checklist') and enumerates the exact payload fields (profile, display_name, focus, required_sections, constraints, escalation_signals). It is unmistakably a local definition read, distinct from advance_qa_orchestration, get_qa_orchestration_catalog, and the other siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the trigger ('Use before reviewing one scoped concern'), an important precondition ('no session is required'), and then splits the two concrete workflows: standalone review via get_qa_orchestration_catalog(), and session primary_review passing the latest current_profile as agent_profile. It also names advance_qa_orchestration as the alternative for applying a session route, so when-to-use and when-not are both explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_qa_orchestrationA

Create a content-free QA session for a tracked review; call once before triage.

Every call creates a separate session. Use prepare_qa_orchestration for a stateless profile checklist. task_type narrows the bundle shortlist; the host still selects the review route.

  • Sessions expire after the configured TTL (1800 seconds by default).

  • Capacity is 100 sessions by default and can be configured.

  • Expired sessions are purged and retained terminal sessions may be evicted. If capacity can only be freed by removing an active session, the call returns a session limit error.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_typeYesBroad request category: ordinary_review for implementation review, widget_review for widget changes, epic_analysis for broad exploration, requirements_analysis for requirement review, qa_planning for coverage planning, autotest_implementation for automation work, or other. This affects recommended_bundles only; the host chooses the route from changed files and evidence; it is not a risk level.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
read_onlyNo
task_typeYes
expires_atYes
next_actionYes
current_stepYes
model_policyNo
allowed_bundlesNo
current_profileNo
deep_assessmentNo
review_profilesNo
selected_bundleNo
allowed_profilesNo
deep_reason_codeNo
selected_profileNo
completed_profilesNo
host_owns_decisionsNo
recommended_bundlesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover safety/idempotency flags; the description adds substantial behavioral context beyond them: TTL expiry (1800s default), 100-session capacity, purge/eviction behavior, and the specific 'session limit' error when only an active session could free capacity. This is exactly the operational detail annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence, then a routing sentence, then tight bulleted constraints. Every sentence carries distinct information with no filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema covers return values, so the description is not obliged to describe them. Combined with annotations and the schema, the description supplies the lifecycle (TTL, capacity, eviction, error) an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum is fully self-documented, so baseline 3 applies. The description still adds real meaning by clarifying what task_type does and does not do ('narrows the bundle shortlist', 'host still selects the review route', 'not a risk level'), correcting a plausible misreading.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create) and resource (content-free QA session) with the timing constraint (call once before triage) front-loaded. It explicitly contrasts with prepare_qa_orchestration, so the agent can separate it from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when ('call once before triage') and when-to-use-the-alternative ('Use prepare_qa_orchestration for a stateless profile checklist'). It also clarifies that every call creates a separate session, preempting accidental reuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.2.0
    • Addeddelete_qa_orchestration
  2. 1 tool updatev0.1.24
    • Addedget_qa_orchestration_catalog
  3. 4 tool updatesv0.1.23
    • Changedadvance_qa_orchestration1 field changed
      • changedInput schema / properties / run_id / description
        Previous value: -"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."New value: +"Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
    • Changedfinish_qa_orchestration1 field changed
      • changedInput schema / properties / run_id / description
        Previous value: -"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."New value: +"Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
    • Changedget_qa_orchestration1 field changed
      • changedInput schema / properties / run_id / description
        Previous value: -"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."New value: +"Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
    • Addedlist_qa_orchestrations

TDQS

A4.9/5.0

Scored across 9 tools

Disambiguation5/5

Each tool targets a distinct action or resource in the QA session lifecycle; even the similarly prefixed get_qa_orchestration* tools are clearly differentiated by their suffixes (session state vs. catalog vs. model policy) and cross-references in descriptions. No overlapping purposes that would cause misselection.

Naming Consistency5/5

All tools use snake_case with a consistent verb_noun pattern around the qa_orchestration object (start, advance, finish, get, list, delete, prepare). The two read tools get_qa_orchestration_catalog and get_qa_orchestration_model_policy extend the pattern predictably by appending the sub-resource.

Tool Count5/5

Nine tools is well-scoped for a QA orchestration server: it covers the full session lifecycle plus static catalog and policy reads without redundant or thin tools. Each tool earns its place.

Completeness5/5

The surface covers complete lifecycle operations: create (start), read (get/list/catalog/policy), update (advance), finalize (finish), and delete, with no obvious dead ends. Recovery of lost IDs and preparation of checklists are also supported.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers