QA Orchestrator
A deterministic MCP service that provides fixed QA review routing, stateful session orchestration, and model policy metadata while the host owns evidence, model calls, and final decisions.
Discover available review profiles and bundles with
get_qa_orchestration_catalog.Get a specialist profile's checklist (focus, required sections, constraints, escalation signals) with
prepare_qa_orchestration.Inspect server-wide provider, model, and reasoning settings for each stage with
get_qa_orchestration_model_policy.Start a content-free QA orchestration session for a task type with
start_qa_orchestration, returning a recommended bundle shortlist.Advance a session through triage, primary review, optional deep review, and synthesis using
advance_qa_orchestration; record completed profiles, risk signals, route selections, or early stops.Read a session's state, next action, and current-stage model policy with
get_qa_orchestration.Recover lost
run_idvalues by listing retained non-expired sessions withlist_qa_orchestrations.Finalize a session with
completed,partial, orblockedoutcome usingfinish_qa_orchestration.Discard one session's memory and optional recovery storage immediately with
delete_qa_orchestration.Does not call models, choose severity/readiness, perform external writes, or store evidence, prompts, source, logs, or model responses.
Allows configuring QA orchestration stages to use OpenAI models, with per-stage model and reasoning policies returned to the host for execution.
QA Orchestrator
QA Orchestrator is a small deterministic FastMCP service for host-owned QA reviews. It keeps orchestration state bounded; evidence, source code, logs, prompts, model responses, and final decisions remain with the primary host agent.
Quick start
Requires Python 3.12+, uv, Git, and a local MCP client with STDIO support.
git clone https://github.com/Rbkmen/qa-orchestrator.git
cd qa-orchestrator
uv sync --locked
uv run qa-orch setup
uv run qa-orchestrator-doctorThen register the server and add the host instructions.
The MCP client starts the server. To inspect the models loaded by that process,
call get_qa_orchestration_model_policy with {} in the client.
See the installation guide for Windows, Claude Code,
installation without a checkout, updates, and troubleshooting.
Related MCP server: Support Ticket Triage MCP
How it works
The primary host obtains authoritative evidence from the required systems and classifies the QA task.
The host calls
start_qa_orchestration. The orchestrator creates a content-free session and returns the first step and its configured policy. Runqa-orch setupto choose OpenAI or Anthropic and configure models and reasoning for each stage.The host runs each stage in its configured model environment and sends the orchestrator only a structured signal after each stage:
triage — select one fixed review bundle or one compatibility profile;
primary review — review every selected profile in the fixed order;
optional deep review — perform one read-only analysis when fixed risk signals match;
final synthesis — consolidate the results. Each stage uses the model and reasoning configured for it in the returned
model_policy. The returned policy selects the provider, model, and reasoning for each stage. Execution speed and latency preferences remain controlled by the user's host/provider settings; the orchestrator does not set or override them.
The host validates findings, runtime evidence, and limitations. For an orchestrated task, it calls
finish_qa_orchestrationwith the samerun_idand its final outcome.
The orchestrator does not call models, choose severity or release readiness, or perform external writes. Its client-rule templates provide portable host-agent instructions for evidence quality and finding presentation; registering the MCP server alone does not load those instructions into the host.
Choose a review path
Use one compatibility profile by default for a routine, narrowly scoped,
low-risk change with one main concern. Use a fixed bundle for broad,
cross-concern, or high-risk changes. For example, a small Ruby guard change can
use ruby_reviewer; a change spanning a Rails endpoint, a background job, and
their tests fits ruby_backend.
recommended_bundles is a task-type shortlist, not a risk score or a required
selection. The host inspects the diff and chooses the initial profile or
bundle. The host sends structured risk signals after primary review; the
orchestrator applies its fixed rules to decide whether optional deep review
follows.
Example: the same review with and without the orchestrator
Ordinary review guided by | Review with QA Orchestrator | |
Small Ruby guard change | The host chooses a reviewer prompt and tracks the review in the conversation. | The host selects |
Broad Ruby change across an endpoint, job, and tests | The host coordinates review steps from its instructions. | The host selects |
Deeper review | The host decides from its own instructions and evidence. | The host sends structured risk signals; the orchestrator applies fixed escalation rules. |
Ownership | The host gathers evidence and makes the final decision. | The host still owns evidence, model calls, findings, and the final decision; the orchestrator receives no source or raw evidence. |
The orchestrator adds a validated workflow contract and bounded progress state. It does not replace host instructions or perform the review itself.
QA Orchestrator at a glance

The diagram shows default memory-only operation. Optional SQLite recovery is described in the MCP interface.
Bundles and profile names
The triage stage selects one fixed bundle or one compatibility profile. start_qa_orchestration returns recommended_bundles as a task-type-based shortlist; it does not restrict allowed_bundles. Choose the route from the changed files and confirmed project stack. The primary-review stage executes bundle profiles sequentially. After every role, the host sends completed_profile, and the orchestrator returns current_profile and completed_profiles. Use the updated session returned by advance_qa_orchestration for the next action and model policy; call get_qa_orchestration only when resuming or recovering a session. Deep review or synthesis is available only after the final role. Transition identifiers are model-neutral; choose the model from the returned model_policy, never from the step name.
Bundle | Profile order |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
autotest and widget include a TypeScript review role; select them when TypeScript review is relevant. Use ordinary_mr for broad non-TypeScript automation, widget_js for broad JavaScript React changes, ruby_backend for broad Ruby backend changes, python_backend for broad Python/MCP changes, and mobile for broad React Native or native iOS/Android changes. Choose from the changed files and confirmed project manifests, not the repository name alone; monorepos can contain several stacks.
Technical profiles and display names:
Profile | Host-facing name |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Faraday is only the internal display name of the code_explorer profile. No external agent, service, package, or model is connected under that name.
Ordinary MR flow with optional escalation:
Triage → Ordinary MR Review
Primary review → Faraday — Evidence Investigator → Code Reviewer → Test Analyzer
├─ no escalation ───────────────────────────────→ Final synthesis
└─ fixed risk signals match → Deep review → Final synthesis
Host → Final QA outcomeWith the final primary-review role, the host must send the structured boolean risk_signals object in the same advance_qa_orchestration call as the final completed_profile (send {} when no signals apply). Deep review is derived only from the fixed signal rules. Do not send risk_signals on the later synthesis transition. The returned session includes matched rules and fixed reason codes as deep_assessment; raw evidence never enters the orchestrator.
Deep-review rules:
high_risk_domain+evidence_uncertain;any two of
cross_system_scope,multiple_plausible_causes,non_reproducible, andhigh_blast_radius;evidence_conflicttogether withhigh_risk_domain,cross_system_scope, orhigh_blast_radius.
For low-risk, narrow reviews, the triage stage may select one compatibility profile instead of a bundle: code_reviewer for a small behavior change, pr_test_analyzer for a test-only change, typescript_reviewer for a TypeScript-only change, react_reviewer for a React-only change, ruby_reviewer for a Ruby-only change, python_reviewer for a Python/MCP-only change, or mobile_reviewer for a React Native/native-platform-only change. Broad or cross-concern reviews continue to use a fixed bundle.
Keep one compact per-task Evidence Packet with stable evidence references (E1, E2, ...) and bounded finding candidates (F-01, F-02, ...). Do not repeat the full diff or raw logs in every model stage.
MCP interface
The service publishes exactly nine tools:
Tool | Purpose |
| Return the fixed checklist for one profile without creating a session |
| Discover all profiles, bundle purposes and ordered routes, and task-type shortlists without creating a session |
| Read the server-wide provider, models, and reasoning for all stages of new sessions; no |
| Create a session and return the task-based bundle shortlist |
| Record one active review step's completion or an early stop; return the next action |
| Read one existing session's step, status, next action, and current-stage model policy |
| Recover lost |
| Finalize the host-owned session outcome after synthesis or a recorded early stop |
| Immediately discard one session's in-memory state and its configured recovery record |
Before starting a session or choosing its triage route, call get_qa_orchestration_catalog with {}. It returns all available profiles with their display names and focus, all bundles with usage guidance and ordered profile IDs, and recommended_bundles_by_task_type. Recommendations are shortlists, not restrictions; choose from changed files and confirmed stack. For a narrow concern, choose one profile and get its detailed checklist through prepare_qa_orchestration. For broad work, select a bundle and preserve its profile order. The catalog is fixed, independent of session/model-policy state, and does not create sessions, change TTL, or write storage.
Bundle orchestration flow:
Triage → Primary review[1] → ... → Primary review[N]
↘ optional Deep review ↗
Final synthesis → Host outcomeSessions are kept in process memory by default. Successful state changes refresh the 1,800-second default TTL; reads do not. The maximum is 100 active sessions, and the shared cache is bounded, so older terminal sessions may be evicted when capacity is needed. Repeating the final call is idempotent while its session is retained. After a restart, the host starts a new session unless optional recovery storage is enabled. read_only=true and host_owns_decisions=true are part of every state.
Set QA_ORCHESTRATOR_SESSION_STORE_PATH to opt into a local SQLite file that restores unfinished orchestration state after a restart. It stores only the current structured session needed for recovery; finalization removes that row. It never stores evidence, prompts, source, logs, model responses, finalized outcomes, history, or statistics. Use one server process per store file. The default remains memory-only.
If a run_id is lost, call list_qa_orchestrations with {}. Its sessions entries include the ID, task type, status, stage, selected route, current profile, and expiry; use these to identify the intended session, then call get_qa_orchestration with its ID. Confirm the intended session if several entries match. The list includes retained terminal sessions and is sorted by expiry descending, with run_id as tie-breaker. It does not change state, renew TTL, or write storage. An empty list means no non-expired sessions are retained by this server; expired or evicted sessions cannot be recovered. After a restart, only unfinished sessions restored from configured recovery storage are available.
After synthesis, the session waits for the host's final outcome. Call finish_qa_orchestration with the session run_id and completed, partial, or blocked. For an early stop, first pass partial or blocked to advance_qa_orchestration, then finish the session with the same outcome. Repeating the same finalization is idempotent; a conflicting outcome is rejected. Tasks that do not use orchestration need no finalization call. The service does not store or report task statistics.
To intentionally discard a session, call delete_qa_orchestration with its run_id. It removes only that session from memory and optional SQLite recovery storage, at any lifecycle stage, and frees its capacity immediately. It returns { "run_id": "qar-...", "deleted": true }; deleted confirms absence, so an already missing or expired ID also succeeds. The ID can no longer be read, advanced, finalized, or restored after restart. Storage failure leaves memory state intact and can be retried. Deletion records no QA outcome and does not stop host tasks or model executions. Use normal early-stop/finalization when the outcome should remain available; delete only when the host intends to discard that session.
Review profiles
prepare_qa_orchestration returns the focus, required sections, constraints, escalation signals, and display name for one profile. For a bundle, the host calls the route for every profile in the returned fixed order and passes its technical identifier in completed_profile after each call.
required_sections is profile-specific: the evidence investigator returns Scope, Evidence Map, and Unverified; test analysis returns Scope, Coverage Gaps, and Unverified; implementation and specialist reviews return Scope, Finding Candidates, Coverage Gaps, and Unverified.
Use the local profile evaluation pack to smoke-check role boundaries and bundle selection without collecting task statistics.
Responsibility boundary
The primary host is responsible for:
obtaining and validating evidence;
calling the issue tracker, code host, test-management system, observability and logging systems, documentation and chat systems, code index, and the file system;
running triage, primary-review, synthesis, and any optional deep-review stage under the policy;
confirmed findings, severity, release/readiness judgment, and the final QA response;
file changes and all external writes.
QA Orchestrator is responsible only for fixed routing, state transitions, read-only constraints, and bounded session finalization. advance_qa_orchestration must not receive an Evidence Packet, prompt, model output, source text, logs, paths, or an arbitrary reason.
The read_only=true flag describes the review boundary. MCP annotations mark
start, advance, and finish as state-changing tools because they update
local orchestration state. Profile lookup, model-policy inspection, and state
lookup are annotated as read-only.
Requirements
Python 3.12+;
uv;an MCP client that supports STDIO;
Windows, macOS, or Linux.
On Windows, uv sync installs the native MCP entry point at
.venv\Scripts\qa-orchestrator-mcp.exe. macOS and Linux use the POSIX
source launcher.
Installation
For the complete setup—including Codex and Claude Code registration, host instructions, verification, and troubleshooting—see the installation guide.
Start with the quick start above to install from a checkout.
On macOS and Linux, the source launcher uses the project's .venv, the active
VIRTUAL_ENV, or an installed qa-orchestrator-mcp from PATH. On Windows,
use the installed .venv\Scripts\qa-orchestrator-mcp.exe entry point. No
separate background process is required.
For clients that support uvx, a checkout is optional:
Replace <commit-sha> with the full commit hash and use the same hash in the
setup and server commands.
uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' qa-orch setupThis saves the model policy in your user configuration. Register the server command with your MCP client so the client starts it when needed; do not run the server command directly in a terminal. For Codex without a checkout:
codex mcp add qa-orchestrator -- uvx \
--from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>' \
qa-orchestrator-mcpUse an immutable commit SHA instead of the default branch for reproducible team configuration. See the installation guide for the remaining no-checkout commands and update instructions.
The setup wizard first lets you choose Russian or English for that run, then
stores only the provider label, model IDs, and the selected provider-specific
reasoning/effort values in the local model-policy.json; the language is not
saved, and the wizard never asks for or stores API keys. Choose a provider and
model that your host client can use; the wizard records the policy but does not
configure provider access. From a checkout, use uv run qa-orch config show and
uv run qa-orch reload. Without a checkout, prefix those commands with
uvx --from 'git+https://github.com/Rbkmen/qa-orchestrator.git@<commit-sha>'.
Colors in the wizard distinguish providers, model IDs, and reasoning values;
set NO_COLOR=1 to disable them or FORCE_COLOR=1 to force them.
The wizard uses a local recommendation/capability catalog and does not contact provider APIs. For custom model IDs, verify that the host account can access the model and supports the selected reasoning/effort value.
Inspect and apply model settings
Check | What it proves |
| The saved policy on disk, or built-in defaults if no file exists |
| The CLI can read and validate that file; the connected server still needs a restart |
MCP | The policy currently loaded by the connected server for new sessions |
Host execution details | Which model actually performed a review stage |
After changing settings, restart the client's server connection and call the MCP policy tool again. The orchestrator provides policy metadata; it cannot switch the host's model or verify that the host executed that model. If the host cannot use a requested model, report that limitation in the review.
Without a saved policy, defaults are gpt-6-luna / max for triage and
gpt-6-sol for primary review / medium, deep review / high, and
synthesis / medium. Setup preserves an existing selection when you accept
its defaults.
Example for Codex with a local checkout:
codex mcp add qa-orchestrator -- "$(pwd)/scripts/qa-orchestrator"On Windows, run this from the repository root in PowerShell:
codex mcp add qa-orchestrator -- "$PWD\.venv\Scripts\qa-orchestrator-mcp.exe"The two supported host integrations are described in the client guides.
Configuration
Variable | Default |
|
|
|
|
|
|
|
|
| unset (disabled) |
The default policy path follows QA_ORCHESTRATOR_DATA_DIR. Configure the same
policy path for setup and for the MCP client; use absolute paths for portable
client configuration. Limits must be positive integers. See the
configuration example for
client environment overrides.
Finalize an orchestration
For completion and early-stop instructions, see the MCP interface section.
Clients and rules
Development
See CONTRIBUTING.md. Any change to the public MCP contract must include an exact tool-surface test and a check that evidence, decisions, and external writes remain with the primary host.
Available Tools
9 toolsadvance_qa_orchestrationRecord QA step resultA
Record one active QA step's result; return updated state and next_action instructions for the host.
For finalization, use finish_qa_orchestration at awaiting_host_outcome; after early stops, match the recorded outcome.
Triage: select exactly one bundle or profile.
Primary review: follow profile order; the final profile requires risk_signals ({} if none).
Early stop: omit route selectors, completed_profile, and risk_signals.
Success saves state and refreshes expires_at.
Invalid/stale requests and unknown/expired run_ids fail without advancing.
Non-idempotent: after a lost response, read get_qa_orchestration(run_id); prior-step replay is rejected.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it. | |
| status | Yes | Use completed after the active step succeeds. Use partial or blocked to stop early; then finalize with the same outcome using finish_qa_orchestration. | |
| risk_signals | No | Required with the final completed primary-review profile. Send an empty object when no evidence-based fixed escalation signals apply; do not send on earlier profiles or on the synthesis transition. | |
| completed_step | Yes | Active step returned in current_step and completed by this call: triage, primary_review, deep_review, or synthesis. When current_step is awaiting_host_outcome, call finish_qa_orchestration instead. | |
| selected_bundle | No | For completed triage, choose one allowed fixed bundle for a broad or cross-concern review. It is mutually exclusive with selected_profile; recommended_bundles is only a shortlist. | |
| selected_profile | No | For completed triage, choose one allowed profile for a narrow, low-risk review. It is mutually exclusive with selected_bundle. | |
| completed_profile | No | After each successful primary-review step, submit exactly the current_profile returned by the previous advance_qa_orchestration call. Omit for other steps or an early stop. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the mutation/non-idempotent profile, but the description goes further: state is saved and expires_at refreshed on success, invalid/stale requests and unknown/expired run_ids fail without advancing, and prior-step replay is rejected. This is exactly the failure-mode and retry context the annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return value in one sentence, then tightly scoped bullets for triage, primary review, early stop, success, failure, and retry. Every line encodes a distinct rule; no filler or repetition of schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, non-idempotent orchestration step with an output schema, the description covers routing, per-step parameter requirements, failure semantics, and recovery, and it correctly does not re-explain return values. An agent has everything needed to call it correctly on the first attempt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds cross-parameter conditional logic the schema does not express: risk_signals is required only with the final completed primary-review profile ({} when none), completed_profile must echo current_profile, and route selectors must be omitted on early stop. It does not restate the enum values themselves, which is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Record one active QA step's result") plus the return contract (updated state and next_action). It explicitly distinguishes itself from the two nearest siblings, finish_qa_orchestration and get_qa_orchestration, so an agent can route without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when/when-not rules and alternatives: use finish_qa_orchestration at awaiting_host_outcome, match the recorded outcome after early stops, do not call on a lost response but read get_qa_orchestration(run_id) instead. Per-step selection rules (triage picks exactly one bundle or profile; early stop omits selectors) leave nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_qa_orchestrationDiscard QA session stateADestructiveIdempotent
Delete one QA session from memory and configured recovery storage. To retain a final QA outcome, use finish_qa_orchestration.
run_id selects one session at any stage; no prior advance or finish required. Recover a lost ID via list_qa_orchestrations(), matching task type, route, and stage; confirm the target if several entries match.
Idempotent: deleted=true confirms absence, including unknown or expired IDs.
Irreversible: get/advance/finish cannot use the ID, even after restart.
Storage failure: memory stays intact; resolve the error and retry.
Host tasks and model executions are unaffected.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| deleted | No | The run_id is absent from this server and its recovery store, including if already absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive, idempotent, and non-read-only, but the description adds material context beyond them: deleted=true confirms absence even for unknown/expired IDs, irreversibility that survives restart, storage-failure behavior where memory stays intact, and the explicit note that host tasks and model executions are unaffected. These are exactly the operational details annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded imperative sentence, then a bulleted block where each line carries a distinct fact (mutation behavior, idempotency, irreversibility, failure recovery, unaffected systems). No sentence restates the name or the annotations verbatim, and there is no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return fields, and it correctly stops there. For a single-parameter destructive mutation with full annotation coverage, everything an agent needs to call it safely and recover from edge cases is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema itself documents the qar- pattern, so the baseline is 3. The description nonetheless adds dispatch-relevant meaning: run_id can select a session at any stage with no prior advance or finish, and it points at list_qa_orchestrations as the recovery path for an unknown ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Delete one QA session from memory and configured recovery storage') and immediately distinguishes itself from the sibling that shares its domain by naming finish_qa_orchestration as the retain-outcome alternative. An agent can differentiate it from advance/get/list/prepare with no schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use this tool ('to retain a final QA outcome, use finish_qa_orchestration' implies deletion is for discarding), relaxes a constraint ('no prior advance or finish required'), and routes the agent to list_qa_orchestrations to recover a lost ID with instructions for disambiguating matches. Explicit alternatives and conditions are provided rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
finish_qa_orchestrationRecord final QA outcomeAIdempotent
Finalize the whole QA run's outcome; active step results belong to advance_qa_orchestration.
After synthesis, use the retained run_id at awaiting_host_outcome to choose completed, partial, or blocked. After an early stop recorded by advance_qa_orchestration, outcome must match its partial/blocked status.
Same outcome: returns the retained terminal session.
Different outcome:
conflicting final outcome; status unchanged.Premature call:
outcome is not ready. Finalized sessions expire at configured TTL (1800s default), or are removed by eviction or restart. Removed run_ids cannot be deduplicated.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it. | |
| outcome | Yes | Host-selected final session status: completed, partial, or blocked. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=false). It discloses idempotent semantics concretely (same outcome returns the retained terminal session; a different outcome raises 'conflicting final outcome' with status unchanged), the TTL expiry (1800s default), eviction/restart removal, and that removed run_ids cannot be deduplicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then bulleted behavioral cases (same outcome / different outcome / premature call) and a closing lifecycle sentence. Dense but every sentence carries distinct information with no repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description instead covers the error paths, state preconditions, and session lifecycle that an agent must know to call this correctly. Nothing material is missing for a two-parameter finalization tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: run_id must be the retained session id from start_qa_orchestration and the outcome choice is constrained by upstream early-stop state. The enum values themselves are already documented in the schema, so it does not fully re-earn a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Finalize the whole QA run's outcome') and immediately scopes it against the sibling that handles the other half of the work ('active step results belong to advance_qa_orchestration'). An agent can distinguish it from advance_qa_orchestration, start_qa_orchestration and the get/list siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions and timing: call after synthesis, use the retained run_id while the session is at awaiting_host_outcome, and ensure the outcome matches a partial/blocked early stop recorded by advance_qa_orchestration. It also warns that a premature call returns 'outcome is not ready', so when-not-to-call is covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_qa_orchestrationInspect one QA sessionARead-onlyIdempotent
Read one retained run_id's content-free state: status, current step, next action, and current-stage model policy.
Use run_id on the server retaining that session to resume or reconcile a lost
advance_qa_orchestration response. Reads do not extend expires_at. Errors unknown run_id or
expired session require a new session via start_qa_orchestration.
If run_id is lost, use list_qa_orchestrations() to find retained sessions first.
For all-stage server policy, use get_qa_orchestration_model_policy(); for review progress or final outcome, use advance_qa_orchestration or finish_qa_orchestration, respectively. Local stdio access requires no additional credentials.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnly/idempotent/closed-world, and the description adds genuinely new behavior: reads do not extend expires_at, the id is server-scoped, and local stdio access needs no extra credentials. These are exactly the operational facts an agent cannot infer from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what is read, then scoping, then error/recovery routing, then disambiguation from siblings. Every sentence carries distinct information; none is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no further explanation, and the description still summarizes the returned state. For a single-parameter read tool, nothing an agent needs to call it correctly or recover from failure is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the pattern plus reuse warning are already in the schema, so baseline is 3. The description adds real meaning by specifying the id must be used on the server retaining that session and that it can be recovered from list_qa_orchestrations, which the schema does not state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (one run_id's retained state) and enumerates exactly what is returned: status, current step, next action, model policy. It is immediately distinguishable from sibling readers like get_qa_orchestration_model_policy and list_qa_orchestrations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly covers when to use it (resume or reconcile a lost advance_qa_orchestration response), what the error cases mean (`unknown run_id`, `expired session`) and the recovery path (start_qa_orchestration), and what to do if the id is lost (list_qa_orchestrations). It also names the alternatives for other needs (model policy, review progress, final outcome).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_qa_orchestration_catalogDiscover QA profiles and bundlesARead-onlyIdempotent
Read the static profile/bundle catalog: available specialists, review scopes, and ordered routes.
Use before creating a session or choosing its triage route; no session or arguments required. Task-type recommendations are shortlists, not restrictions. Choose a bundle for broad work or a single profile for a narrow concern using changed files and confirmed stack.
Detailed checklist: prepare_qa_orchestration(agent_profile).
Create a session: start_qa_orchestration; set its route: advance_qa_orchestration.
Session IDs: list_qa_orchestrations().
Available routes come from the installed server's bundled definitions; active sessions and configured models do not change this catalog.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| bundles | Yes | |
| profiles | Yes | |
| read_only | No | |
| host_owns_decisions | No | |
| recommended_bundles_by_task_type | Yes | Task-based shortlists only; they do not restrict the available bundles. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false. The description adds useful context beyond them: the catalog is static and comes from the installed server's bundled definitions, active sessions and configured models do not change it, and task-type recommendations are shortlists rather than restrictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose, then moves into usage conditions, cross-references, and a final caveat. The bullet list for related tools is efficient and nothing reads as filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only discovery tool with an output schema, the description provides everything an agent needs: what it returns, when to call it, what preconditions are unnecessary, and how it relates to sibling tools. Return-value details are appropriately left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the rubric baseline is 4. The description adds a small amount of useful clarity by stating that no session or arguments are required, though there is no parameter syntax or structure to explain further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: reading the static profile/bundle catalog, with its contents enumerated (specialists, review scopes, ordered routes). It is clearly distinguishable from siblings like prepare_qa_orchestration and start_qa_orchestration because it describes discovery, not session mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this before creating a session or choosing its triage route, and notes that no session or arguments are required. It also provides next-step alternatives for detailed checklists, session creation, route setting, and listing session IDs, and explains how to choose a bundle versus a single profile.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_qa_orchestration_model_policyInspect server model settingsARead-onlyIdempotent
Read server-wide model settings for every stage: provider, model IDs, and reasoning.
Use to inspect the policy loaded for new sessions; takes no arguments.
Existing session state and current-stage policy: get_qa_orchestration(run_id).
Create a session with these settings: start_qa_orchestration.
Reads the loaded in-memory snapshot, not the current policy file; restart the server connection after file changes. No session is created and no provider is contacted; the result does not verify which model the host executed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| provider | Yes | |
| deep_model | Yes | |
| triage_model | Yes | |
| primary_model | Yes | |
| deep_reasoning | No | |
| synthesis_model | Yes | |
| triage_reasoning | No | |
| primary_reasoning | No | |
| synthesis_reasoning | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and closed-world behavior, but the description adds substantial non-obvious context: it reads an in-memory snapshot rather than the current policy file, requires a server restart after file changes, creates no session, contacts no provider, and does not verify which model the host executed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and scope, then concisely groups usage routing into bullets and behavioral caveats into tight sentences. Every sentence adds distinct value with no repetition or wasted space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need not be described. The description covers what is read, when to use it, alternatives, and important behavioral caveats, making it complete for a zero-argument inspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description confirms 'takes no arguments,' which is consistent with the empty schema and leaves no parameter ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: reads server-wide model settings for every stage, listing provider, model IDs, and reasoning. It also distinguishes itself from siblings by explicitly naming get_qa_orchestration(run_id) and start_qa_orchestration as alternative tools for different needs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it to inspect the policy loaded for new sessions and notes it takes no arguments. It routes the agent clearly: existing session/current-stage policy uses get_qa_orchestration, and creating a session with these settings uses start_qa_orchestration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_qa_orchestrationsList retained QA sessionsARead-onlyIdempotent
List non-expired QA session identifiers and brief metadata retained by this server.
Recover a lost run_id by matching task type, route, and stage; confirm if several match. Call get_qa_orchestration(run_id) for full state and next action, or get_qa_orchestration_model_policy() for server policy. Takes no arguments.
Includes retained terminal sessions; expired or evicted sessions are unrecoverable.
Sorts by expires_at descending, then run_id descending; empty means none are retained.
After restart, only unfinished sessions restored from configured local storage appear.
Reading does not extend session TTL.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| sessions | Yes | |
| read_only | No | |
| host_owns_decisions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/openWorldHint=false, so the safety profile is covered; the description goes further with non-obvious behavior the schema cannot express: reading does not extend TTL, expired/evicted sessions are unrecoverable, sorting is expires_at desc then run_id desc, empty result means nothing retained, and post-restart only unfinished local-storage sessions reappear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence states purpose, the second sentence front-loads the primary use case and alternatives, and four tight bullets carry only non-redundant behavioral facts (TTL, sort, restart, empty semantics). No sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists (so return-field documentation is not required), the description still supplies everything an agent needs that the structured fields do not: empty-result interpretation, sort order, TTL non-extension, and restart scoping. Nothing material to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and schema description coverage is 100%, so the baseline is 4. The description explicitly confirms "Takes no arguments," which is the only parameter-level fact available and is already fully specified by the schema, so it cannot exceed the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List non-expired QA session identifiers and brief metadata") and immediately distinguishes itself from get_qa_orchestration, which returns "full state" rather than a list. An agent can tell from the description alone which of the siblings to call for a lookup versus a state read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ("Recover a lost run_id by matching task type, route, and stage; confirm if several match") and routes to the correct alternatives by name for other needs (get_qa_orchestration for full state, get_qa_orchestration_model_policy for server policy). It also states a when-not: expired or evicted sessions are unrecoverable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_qa_orchestrationRead specialist review checklistARead-onlyIdempotent
Return one specialist's fixed ReviewRoute checklist: profile, display_name, focus, required_sections, constraints, and escalation_signals.
Use before reviewing one scoped concern; no session is required.
Standalone review: use get_qa_orchestration_catalog() to match profile focus to changed files and confirmed stack.
Session primary_review: pass the latest current_profile as agent_profile; preserve review_profiles order for bundles.
This local lookup reads bundled definitions, without executing reviews or calling models. Use advance_qa_orchestration to apply a session route. Results depend only on agent_profile. Invalid values fail input validation; retry with a listed enum value.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_profile | Yes | Choose one fixed specialist profile: code_explorer maps execution and evidence; code_reviewer reviews implementation; pr_test_analyzer reviews coverage; security_reviewer and silent_failure_hunter target security and silent failures; ruby_reviewer, python_reviewer, typescript_reviewer, react_reviewer, and mobile_reviewer target their named stacks. Choose from changed files and confirmed stack. |
Output Schema
| Name | Required | Description |
|---|---|---|
| focus | Yes | The profile's review boundary and primary responsibility. |
| profile | Yes | |
| read_only | No | |
| constraints | Yes | Shared limits for this read-only review profile. |
| display_name | Yes | |
| required_sections | Yes | Profile-specific output sections; use these instead of a generic template. |
| escalation_signals | No | Conditions to assess against evidence. A listed condition is not itself a deep-review trigger; set risk_signals only when the corresponding evidence-based condition applies. |
| host_owns_decisions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint/openWorldHint/idempotentHint annotations, it discloses that this is a local lookup over bundled definitions that executes no reviews and calls no models, that results are a pure function of agent_profile, and that invalid values fail input validation and should be retried with a listed enum value. This error-handling and no-side-effect disclosure is genuinely non-obvious context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by tightly scoped bullets for the two usage paths and a short closing note on local behavior and validation. Every sentence carries information; nothing is redundant with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one required enum parameter, a 100%-covered schema, and an output schema present, the remaining burden is routing and behavior – both covered (catalog vs. session path, advance_qa_orchestration alternative, determinism, validation failure). Nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum description already defines all ten profiles, so the schema carries the baseline. The description still adds selection guidance not in the schema: results depend only on agent_profile, choose from changed files and confirmed stack, and preserve review_profiles order when handling bundles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb and resource ('Return one specialist's fixed ReviewRoute checklist') and enumerates the exact payload fields (profile, display_name, focus, required_sections, constraints, escalation_signals). It is unmistakably a local definition read, distinct from advance_qa_orchestration, get_qa_orchestration_catalog, and the other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the trigger ('Use before reviewing one scoped concern'), an important precondition ('no session is required'), and then splits the two concrete workflows: standalone review via get_qa_orchestration_catalog(), and session primary_review passing the latest current_profile as agent_profile. It also names advance_qa_orchestration as the alternative for applying a session route, so when-to-use and when-not are both explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_qa_orchestrationA
Create a content-free QA session for a tracked review; call once before triage.
Every call creates a separate session. Use prepare_qa_orchestration for a stateless profile
checklist. task_type narrows the bundle shortlist; the host still selects the review route.
Sessions expire after the configured TTL (1800 seconds by default).
Capacity is 100 sessions by default and can be configured.
Expired sessions are purged and retained terminal sessions may be evicted. If capacity can only be freed by removing an active session, the call returns a
session limiterror.
| Name | Required | Description | Default |
|---|---|---|---|
| task_type | Yes | Broad request category: ordinary_review for implementation review, widget_review for widget changes, epic_analysis for broad exploration, requirements_analysis for requirement review, qa_planning for coverage planning, autotest_implementation for automation work, or other. This affects recommended_bundles only; the host chooses the route from changed files and evidence; it is not a risk level. |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| status | Yes | |
| read_only | No | |
| task_type | Yes | |
| expires_at | Yes | |
| next_action | Yes | |
| current_step | Yes | |
| model_policy | No | |
| allowed_bundles | No | |
| current_profile | No | |
| deep_assessment | No | |
| review_profiles | No | |
| selected_bundle | No | |
| allowed_profiles | No | |
| deep_reason_code | No | |
| selected_profile | No | |
| completed_profiles | No | |
| host_owns_decisions | No | |
| recommended_bundles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover safety/idempotency flags; the description adds substantial behavioral context beyond them: TTL expiry (1800s default), 100-session capacity, purge/eviction behavior, and the specific 'session limit' error when only an active session could free capacity. This is exactly the operational detail annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then a routing sentence, then tight bulleted constraints. Every sentence carries distinct information with no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema covers return values, so the description is not obliged to describe them. Combined with annotations and the schema, the description supplies the lifecycle (TTL, capacity, eviction, error) an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum is fully self-documented, so baseline 3 applies. The description still adds real meaning by clarifying what task_type does and does not do ('narrows the bundle shortlist', 'host still selects the review route', 'not a risk level'), correcting a plausible misreading.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (content-free QA session) with the timing constraint (call once before triage) front-loaded. It explicitly contrasts with prepare_qa_orchestration, so the agent can separate it from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when ('call once before triage') and when-to-use-the-alternative ('Use prepare_qa_orchestration for a stateless profile checklist'). It also clarifies that every call creates a separate session, preempting accidental reuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.2.0- Added
delete_qa_orchestration
1 tool update
v0.1.24- Added
get_qa_orchestration_catalog
4 tool updates
v0.1.23- Changed
advance_qa_orchestration1 field changed- changed
Input schema / properties / run_id / descriptionPrevious value: -"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."New value: +"Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
- Changed
finish_qa_orchestration1 field changed- changed
Input schema / properties / run_id / descriptionPrevious value: -"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."New value: +"Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
- Changed
get_qa_orchestration1 field changed- changed
Input schema / properties / run_id / descriptionPrevious value: -"Session identifier returned by start_qa_orchestration: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."New value: +"Session identifier returned by start_qa_orchestration or list_qa_orchestrations: qar- followed by 32 lowercase hexadecimal characters. Reuse it for this session; do not invent or transform it."
- Added
list_qa_orchestrations
TDQS
Scored across 9 tools
Each tool targets a distinct action or resource in the QA session lifecycle; even the similarly prefixed get_qa_orchestration* tools are clearly differentiated by their suffixes (session state vs. catalog vs. model policy) and cross-references in descriptions. No overlapping purposes that would cause misselection.
All tools use snake_case with a consistent verb_noun pattern around the qa_orchestration object (start, advance, finish, get, list, delete, prepare). The two read tools get_qa_orchestration_catalog and get_qa_orchestration_model_policy extend the pattern predictably by appending the sub-resource.
Nine tools is well-scoped for a QA orchestration server: it covers the full session lifecycle plus static catalog and policy reads without redundant or thin tools. Each tool earns its place.
The surface covers complete lifecycle operations: create (start), read (get/list/catalog/policy), update (advance), finalize (finish), and delete, with no obvious dead ends. Recovery of lost IDs and preparation of checklists are also supported.
Maintenance
Related MCP Connectors
Hosted MCP for denial, prior auth, reimbursement, workflow validation, batch scoring, and feedback.
The MCP server that vets MCP servers: identity, risk grade and per-tool risk before you install.
Hybrid human + AI expertise for faster, trusted answers and decisions via MCP Server.
A paid remote MCP for hosted MCP server, built to return verdicts, receipts, usage logs, and audit-r
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server for healthcare claims workflow scoring, validation, and feedback, supporting denial risk, prior authorization, and reimbursement assessment.838 PyPIMIT
- FlicenseAqualityFmaintenanceA local MCP server for governed support-ticket triage that reads synthetic tickets and knowledge articles, prepares evidence-backed recommendations, and records local audit events.9-
- AlicenseCqualityDmaintenanceAn MCP server that guides QA and verification processes by breaking down tasks into manageable steps and providing LLM-driven, confidence-scored tool recommendations.112 npm6MIT
- AlicenseAqualityCmaintenanceA self-hosted MCP server offering provider-independent AI routing, deterministic verification, and cryptographically signed, tamper-evident audit trails for every response.831 npm2MIT