Skip to main content
Glama

llm-migrate

Migrate your LLM application to a new model, provider, or platform — for example Anthropic Claude ↔ OpenAI GPT, or a direct API ↔ Amazon Bedrock — with agent-assisted research, a reviewable migration plan, evidence-linked adaptations you accept or reject change by change, and before/after evaluation, all before changing production code.

CI

llm-migrate helps a developer or coding agent inspect an existing Python application, research the source and target models, and decide exactly what a safe migration requires — without editing any application file itself.

Use it when you are:

  • upgrading to a newer model generation

  • moving between providers, such as Anthropic and OpenAI

  • moving between platforms, such as a direct API and Amazon Bedrock

  • replacing a deprecated model

  • comparing targets for capability, lifecycle, cost, or latency

  • validating prompt, tool, structured-output, or invocation changes

The project is local-first and provider-neutral. It does not silently rewrite application files, operate a hosted inference gateway, or promote researched facts into the shared registry without review.

The primary experience is an agent host connected to the llm-migrate MCP server, driving one guided migration run. Codex, Claude Code, or another MCP-capable host supplies the generative models and web/search tools. llm-migrate supplies the scanner, typed research stages, evidence gates, session registry, migration planner, per-file adaptation worklist, per-change review, and evaluation contracts.

One start_migration call creates a run workspace, matches both model identifiers against the reviewed registry (tolerating vague spellings), scans the application, decides whether research is needed and says why, and returns ordered next steps. From there the run is a state machine — get_run_status reports the position and the single next action at any point:

  1. Optionally research missing or stale facts with host-supplied agents, independently review every consequential claim, and build an expiring, user-scoped session registry.

  2. Work the per-file adaptation worklist: the agent writes adapted prompts and files and submits them with guidance dispositions and evidence-linked annotated changes; unaffected files close in one confirmation call.

  3. If the plan has blockers, each becomes a question for you with registry-backed options; the agent presents them verbatim and records your decisions durably.

  4. Review every annotated change like a pull request — accept or reject each one; rejections regenerate the deliverable deterministically.

  5. Finalize, validate, and finalize again. The first finalize checks the whole deliverable set together and writes the manifest, the report (which opens with an "Action required" list), the change log, and a generated request-shape contract test. You run that test or a BYOK evaluation, the agent records the outcome, and a second finalize records the migration as validated.

If the canonical registry already has fresh coverage, the research request is refused and the agent continues with the reviewed local knowledge. "Live research by default" therefore means always check and research when needed, not "browse even when verified facts already exist."

Related MCP server: FreshStack MCP

Agentic workflow

flowchart TD
    U["Developer + application repository"] --> H["Agent host<br/>Codex, Claude Code, or custom host"]
    H --> M["llm-migrate MCP server"]
    M --> S["start_migration<br/>match models + scan + decide research"]
    S --> K{"Registry knowledge<br/>complete and fresh?"}

    K -- Yes --> W["Per-file adaptation worklist"]
    K -- No --> R["Bounded research request +<br/>generated researcher/reviewer prompts"]
    R --> A["Host-supplied research agents<br/>generative model + web/search"]
    A --> V["Independent evidence reviewer<br/>refetch cited sources"]
    V --> C["Deterministic validation + consensus"]
    C --> O["Immutable, expiring<br/>session registry overlay"]
    O --> W

    W --> B{"Blockers?"}
    B -- Yes --> D["Questions with registry-backed options<br/>presented verbatim; user decides"]
    D --> W
    B -- No --> T["Agent submits adapted prompts/files<br/>with evidence-linked annotated changes"]
    T --> Y["Per-change review<br/>user accepts or rejects each change"]
    Y --> Z["Validation disposition +<br/>consistency-gated finalize"]
    Z --> E{"Run evaluation?"}
    E -- Yes --> X["Source + target evaluation"]
    X --> G["Regression analysis +<br/>bounded optimization"]
    G --> Q["Human review"]
    E -- No --> Q
    Q --> I["Implement migration outside llm-migrate"]

The agent host makes generative calls and writes the adaptations. The MCP server remains the deterministic control and validation layer: it derives worklists, validates fail-closed, stores deliverables under the run's output/, and never modifies the application tree.

Prerequisites

For the recommended guided workflow:

  • Python 3.11 or newer

  • Git

  • a local Python application repository

  • an MCP-capable coding agent with a generative model — use the most capable reasoning model your host offers, with extended thinking enabled (for example an Opus- or Fable-class Claude model, or a GPT-5-class model at high reasoning effort). The agent researches, adapts prompts and files with evidence-linked annotations, and relays review decisions; small or fast tiers produce more rejected submissions and retries

  • web/search access in that agent host (for research on demand)

  • a separate agent or fresh isolated context for evidence review

  • the source and target model identifiers as you know them — vague spellings are matched registry-first, and anything ambiguous returns candidates for you to confirm

Provider credentials are not needed for scanning, research validation, planning, adaptation, or reporting. They are only needed when you explicitly run source/target model evaluations. Amazon Bedrock evaluation also requires the optional aws extra and your normal local AWS configuration.

Installation

Easiest: let your coding agent install it

If your coding agent can run shell commands and edit its own MCP configuration (Claude Code, GitHub Copilot, Codex, Cursor, ...), paste this prompt and let it do the setup:

Install the llm-migrate MCP server for me:
1. git clone https://github.com/Athenaxlee/llm-migrate.git into a tools
   directory of your choice (tell me where), or reuse an existing checkout.
2. Inside the checkout, create a Python 3.11+ virtual environment at .venv and
   run: python -m pip install --upgrade pip && python -m pip install .
   (use '.[aws]' instead of '.' if I plan to run Amazon Bedrock evaluations).
3. Verify it works: .venv/bin/llm-migrate registry validate
   (on Windows: .venv\Scripts\llm-migrate registry validate).
4. Register the MCP server in THIS host's own MCP configuration, pointing the
   command at the absolute path of .venv/bin/llm-migrate-mcp
   (Windows: .venv\Scripts\llm-migrate-mcp.exe). Do not change any other
   configuration.
5. Show me the config change and the verification output, and tell me to
   restart/reload so the server is picked up.

For example, on Claude Code the registration step is:

claude mcp add llm-migrate -- /absolute/path/to/llm-migrate/.venv/bin/llm-migrate-mcp

Manual install

git clone https://github.com/Athenaxlee/llm-migrate.git
cd llm-migrate

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install .

Windows PowerShell activation:

.venv\Scripts\Activate.ps1

Optional Amazon Bedrock support:

python -m pip install '.[aws]'

Verify the installation:

llm-migrate registry validate
llm-migrate models list

Updating an existing install

When this repository gets new commits and the MCP server is already installed, update the same checkout in place. The MCP configuration keeps pointing at the same .venv executable, so no configuration change is needed:

cd /path/to/llm-migrate              # the checkout your MCP config points at
git pull
.venv/bin/python -m pip install .    # Windows: .venv\Scripts\python -m pip install .
.venv/bin/pip show llm-migrate       # confirm the new version
.venv/bin/llm-migrate registry validate

Then restart or reload your MCP client (or just that server entry): hosts keep the stdio server process running and will not pick up new code until the server restarts. A development install (pip install -e '.[dev]') only needs the git pull and the restart.

Or paste this prompt and let your coding agent do it:

Update my llm-migrate MCP server: find the checkout my MCP configuration points
at, run git pull there, reinstall it into that checkout's .venv with
"python -m pip install .", verify with "llm-migrate registry validate", and
show me the installed version from "pip show llm-migrate". Do not change the
MCP configuration unless the executable path actually moved. Then tell me to
restart/reload the MCP server so the update takes effect.

Connect the MCP server

The installed stdio server is:

llm-migrate-mcp

Each agent host keeps its own MCP registry. Registering the server in one host does not make it available in another — a server added to VS Code's mcp.json serves GitHub Copilot there but is invisible to Claude Code, and vice versa. Register it in every host you use, always pointing at the executable inside the virtual environment:

Host

How to register

Claude Code

claude mcp add llm-migrate -- /absolute/path/to/llm-migrate/.venv/bin/llm-migrate-mcp

VS Code (GitHub Copilot)

add the server to .vscode/mcp.json (workspace) or your user mcp.json

Codex and other hosts

add the server to that host's MCP configuration file

A generic configuration block (the filename and wrapper syntax vary by host; on Windows the executable is .venv\Scripts\llm-migrate-mcp.exe):

{
  "mcpServers": {
    "llm-migrate": {
      "command": "/absolute/path/to/llm-migrate/.venv/bin/llm-migrate-mcp"
    }
  }
}

Verify in each host after restarting or reloading it: the host's tool list should include start_migration (in Claude Code, claude mcp list shows the server as connected). If the agent says it has no llm-migrate tools, the server is registered in a different host — or the host was not restarted after registration.

The agent host must provide its own generative model and web/search capability. The llm-migrate MCP server does not contain an embedded model or general web search tool, and it never calls an LLM API itself: research, prompt adaptation, and review are done by your coding agent with its own model. Pick that model accordingly (see Prerequisites): the guided workflow is demanding, and a stronger reasoning model with extended thinking finishes it in fewer turns.

Hosts with tight inline-tool budgets can set the environment variable LLM_MIGRATE_TOOLSET=guided on the server process to expose only the 23 guided-workflow tools; the default (full) exposes everything. Production migrations can pass strict=true to start_migration (CLI run start --strict) so unknown evidence, missing invocation facts, incomplete coverage, consistency findings, and a missing validation disposition block delivery instead of warning.

First migration with an agent

Open the application repository in your MCP-capable agent and start the guided run. A useful starting instruction is:

Use the llm-migrate MCP tools to migrate this application.

Application: /absolute/path/to/application
Source: <platform> / <model id as you know it>
Target: <platform> / <model id as you know it>

Start with start_migration and follow its next_steps. When research is
recommended, run it without asking me, with non-interactive background
agents — it needs nothing from me. Ask me before finalizing. Present every
blocker question and review change to me verbatim — do not decide on my
behalf.

Research, when the registry's recorded facts are missing or past their freshness window, runs one researcher and one independent reviewer agent per scope. The generated prompts tell those agents to use web search and page-fetch tools and never a visible browser, and the orchestration guidance runs scopes in parallel, so the stage needs no attention from you. To skip it for a quick pass, add research="skip" to start_migration (CLI --skip-research).

For production migrations, add strict=true: unknown evidence URLs are rejected, missing invocation facts and incomplete prompt coverage block, and finalize refuses to complete cleanly over gaps, findings, undecided changes, or a missing validation disposition.

The run workspace defaults to <application>/.llm-migrate/runs/<run-id>/ (pass output_dir to relocate it) and collects:

Artifact

Purpose

migration.yaml

The run's durable identity: models, platforms, endpoints, and the selector-qualified invocation ids

request.yaml

Bounded research request, written only when knowledge is missing or stale, with the reasons

Research and review artifacts

Source-backed claims plus independent claim-level verdicts, validated in place

Session registry

Expiring, hash-linked, visibly non-canonical knowledge for this migration

decisions.yaml

Your recorded blocker decisions, re-applied on every plan regeneration

output/prompts/, output/files/

Validated adapted deliverables — the application tree is never modified

changes.yaml

Every submission's rationale, guidance dispositions, and annotated changes with evidence

change-decisions.yaml

Your per-change accept/reject decisions

output/migration-manifest.yaml

The machine-readable plan: changes, blockers, warnings, unknowns, tests, rollout

output/migration-report.md

Human-readable report with per-file changes, evidence, and the decision trail

output/validation/test_target_invocation.py

Generated request-shape contract test you wire up and run yourself

Key behaviors during the run:

  • Model identifiers are matched registry-first, tolerating Bedrock cross-region prefixes (us.anthropic...), version suffixes (-v1:0), spacing/typos, and loose platform names ("bedrock"). Where the reviewed registry says the target's bare model id is not invocable on demand (for example Claude on Bedrock), start returns the reviewed invocation selectors for you to confirm, and every later surface enforces the selector-qualified id.

  • Every prompt task's verbatim_source is the unmodified original, never a proposed adaptation. Submissions must dispose every guidance item (applied / not applicable / declined with a note) and document every edit as an anchored, evidence-linked annotated change, reconciled against the real diff of the decoded runtime values — undocumented or phantom edits are rejected, new files need at least one evidence-linked annotation, and a file that genuinely needs no change is recorded with unchanged=true. Files that only import an SDK incidentally close through one confirm_unaffected call, and submit_adaptations batches submissions in one locked write.

  • Blockers are never dead ends: each yields the question to ask you plus registry-backed options with consequences and evidence URLs — retarget, redesign, correction, or an explicit accept that requires your own rationale and is never a default. Decisions whose blocker disappears are reported stale, never silently applied.

  • Finalize checks the whole effective deliverable set together: lingering source references (including reviewed aliases of the source model), forbidden bare target ids, mixed selectors, missing target attribution, and undisposed dropped couplings. It also sweeps every scanned application file: a file that still names the source model with no deliverable or recorded review is reported, whatever the scanner failed to trace. A mention you want to keep (documentation, history) is acknowledged with confirm_unaffected(acknowledge_source_references=true) and your own rationale.

  • Validation must be evidenced: generated_tests needs your passing test result and its summary line, and byok_evaluation needs the evaluation run artifact bound to the current manifest. An evaluation of an older plan, or a record without evidence, is reported NOT VALIDATED.

  • Prompt discovery survives real repository layouts: config paths written relative to the repository root resolve when you scan a subdirectory, Windows-style backslash values and case differences resolve, and simple open(...).read() / Path(...).read_text() loaders are traced. When prompt consumers exist but no prompt file resolved, start_migration lists the candidate files for you to confirm before writing anything; on a live run, add_prompt_sources and confirm_prompt_consumer fix discovery in place instead of forcing a new run.

  • Unknowns give directions: each one says why it matters for your application, the exact next call (with the run directory filled in), and what closes it. Where the reviewed registry records conflicting evidence (for example whether adaptive thinking can be disabled on Bedrock), the run emits a small probe script under output/probes/ for you to run with your own credentials; its printed result is recorded with record_observation for this run only and never changes the registry.

  • Review has teeth: adapted pricing values are checked against the registry's facts unit-aware (per token / per 1K / per 1M) — a price that departs from the target's recorded price is flagged with its implied factor unless the change cites registry-recorded evidence, and leftover source pricing is flagged stale. A change that depends on a contested registry fact is marked CONTESTED in change review until an observation resolves it, and after the first finalize the run points you at a drafted evaluation instead of letting validation be skipped.

  • Research runs by default and hands-free: when facts are missing or past their freshness window, next_steps and get_run_status tell the host to run the research stages now with non-interactive background agents (web search and page-fetch tools, never a visible browser), scopes in parallel, and to ask you first only if you asked to be consulted. A fresh registry (see "Keeping the registry fresh") means no research at all.

  • Noise stays out of your way: prompt files referenced only by another model's configuration profile (or unreferenced next to selected siblings) are reported out of scope, unknowns the scan already answers are not raised, and registry guidance whose trigger (structured output, tool use, reasoning) your application never exhibits is pre-marked not applicable with the reason.

MCP tool guide

The tools are grouped in the order an agent normally uses them. Most users drive everything through the guided workflow group; the rest remain available for manual composition.

Tool

Use it when

What it does

start_migration

Begin any migration

Matches both models registry-first, scans, decides whether research is needed, writes the run workspace, returns next steps

get_run_status

Any time

Reports the run's state-machine position and the single next action

get_research_prompts

The run recommends research

Returns scope-isolated researcher/reviewer prompt pairs with exact bounds, schemas, and artifact paths; the prompts demand non-interactive tools and the guidance runs scopes in parallel

validate_research_artifact

A research/review artifact is written

Validates the YAML in place with the deterministic gates

build_session_registry

All required scope artifacts validate

Finalizes the immutable, expiring session overlay

list_adaptation_tasks

The plan is ready

Derives the per-file worklist (snapshot-backed) with evidence-linked guidance and unaffected_files

get_blocker_resolutions / record_blocker_decision

The plan has blockers

Presents each blocker's question and registry-backed options; records your durable decision

submit_adapted_prompt / submit_adapted_file / submit_adaptations

The agent wrote adaptations

Validates fail-closed (dispositions, annotated changes, decoded-value diffs, selector enforcement) and stores deliverables; submit_adaptations batches

confirm_unaffected

Incidental-SDK files remain

Records reviewed no-change entries for them in one call

scaffold_evaluation

Status reports validation_pending after the first finalize

Drafts DRAFT evaluation cases from your application's own sample inputs and a suite bound to the current plan, for you to review and run

record_observation

You ran a probe or check for an open unknown

Records what the target actually did as a run-scoped observation that closes the unknown; the registry is never changed

add_prompt_sources / confirm_prompt_consumer

Status reports discovery_incomplete

Adds prompt files to the live run, confirms which file a dynamic consumer reads, or records your dismissal with its rationale; the worklist re-derives in place

get_change_review / record_change_decision / record_change_decisions

Deliverables await review

Presents each annotated change verbatim; your accept/reject regenerates the deliverable deterministically

record_validation_disposition

After the first finalize

Records how the migration was validated (a passing test result, or an evaluation run bound to the current manifest)

finalize_migration

Everything is decided

Runs the cross-surface consistency gate and writes the manifest, report, change log, and contract test

resolve_model / get_model_profile

Identity questions during the run

Registry-first resolution and full reviewed profiles

2. Understand the application and models

Tool

Use it when

What it does

scan_application

Start any repository analysis

Scans a local Python application (and its YAML/JSON/TOML prompt configuration) into a normalized, source-located coupling inventory without executing it

resolve_model

You have a name, alias, or platform model ID

Resolves it to one canonical model and optional platform representation; ambiguity returns ranked candidates instead of guessing

get_model_profile

You need all reviewed facts for one known model

Returns the validated local registry profile and provenance

check_model_lifecycle

Deprecation or end-of-life may drive the migration

Interprets reviewed lifecycle facts at a selected date without live research

compare_models

Source and target are known

Reports same, different, unsupported, and unknown states across the migration surface, each claim evidence-linked

recommend_models

The target is not yet chosen

Hard-filters incompatible registry models, then ranks the remaining candidates against application requirements and goals

query_live_pricing

You explicitly want a current OpenRouter price observation

Fetches time-stamped third-party pricing evidence without changing canonical facts

estimate_migration_cost

You know the workload shape

Estimates recurring source and target token cost from checked-in canonical pricing

Because the registry is intentionally small, recommend_models only ranks models it knows. Use the research tools when the intended target is missing or its migration-critical facts are stale.

3. Research missing or stale knowledge

Tool

Use it when

What it does

create_migration_research_request

Begin a researched workflow outside a guided run

Scans locally and creates a bounded request only for missing or stale topics; refuses unnecessary research

validate_research_result

A research agent has returned a typed artifact

Checks schema, scope, source policy, references, identities, and requested-topic boundaries before review

validate_evidence_review

An independent reviewer has returned verdicts

Confirms reviewer independence, research hash linkage, and claim coverage

validate_research_artifact

Artifacts live in a run workspace

Validates researcher/reviewer YAML in place by path and scope

build_research_consensus

Research and independent review both validate

Deterministically accepts or rejects claims; agent agreement alone is never evidence

build_session_registry

All required scope artifacts are present in the run workspace

Finalizes an immutable, expiring overlay; missing work or high-impact conflicts fail closed

generate_session_migration_plan

The target depends on session knowledge

Generates a plan over canonical plus explicitly selected session knowledge and exposes its trust and expiry

propose_registry_update

Researched facts should be considered for the shared registry

Produces a review-only canonical update proposal; it never edits or promotes registry files

Live discovery happens in the agent host, not inside these tools. Research agents receive the bounded request and use the host's generative model and search tools. The reviewer must independently refetch cited sources. See the agent-host workflow for the artifact protocol.

4. Prepare the migration (low-level)

Tool

Use it when

What it does

analyze_prompt

You need to understand one prompt's intent and assumptions

Conservatively identifies objectives, contracts, instructions, examples, grounding, reasoning, tool, and verbosity characteristics without inference

prepare_prompt_migration

You need a reviewable prompt candidate

Applies deterministic, registry-backed mappings and returns the candidate, semantic diff, risks, and validation guidance without rewriting the source file

validate_prompt

You want target-specific static checks

Checks context capacity, tool/output needs, reasoning instructions, and target platform compatibility

analyze_invocation

You need the SDK/request/tool/output coupling for an application

Normalizes provider operations, parameters, tool schemas, output contracts, streaming, reasoning controls, and multimodal payloads

prepare_invocation_migration

Source and target endpoints are known

Produces the target SDK, operation, model ID, parameters, tool/output candidates, configuration, warnings, and blockers without editing code

generate_migration_plan

Canonical registry knowledge is sufficient

Composes scanning, comparison, prompt/invocation preparation, validation, tests, and rollout into one application-level plan

generate_migration_report

A person needs to review the plan

Renders the integrated migration workflow as a readable Markdown report

Prompt preparation itself does not call a generative model: the deterministic candidate is the source prompt verbatim, with evidence-linked guidance. The model-authored rewrite belongs to the agent host, and the guided workflow then holds it to dispositions, annotations, and per-change review. That rewrite is deliberately not a hidden core operation.

5. Evaluate and improve

Tool

Use it when

What it does

generate_eval_suite

You have a migration plan and representative cases

Binds one immutable evaluation corpus to the migration-manifest hash

run_migration_eval

Source and target endpoint configs are ready

Runs the same suite against both endpoints using user-owned credentials

compare_outputs

You need a deterministic comparison for one case

Compares paired results, including quality, latency, tokens, cost, refusal, errors, tools, and structured output

analyze_regressions

An evaluation run is complete

Produces a categorized regression report without inventing unsupported diagnoses

optimize_migration

You have regression evidence and optional candidate runs

Returns bounded, reproducible, review-only recommendations under quality, cost, latency, and run limits

Endpoint configurations name credential environment variables; they never contain credential values. Built-in execution supports direct OpenAI, direct Anthropic, and Amazon Bedrock Converse. Other platforms require an executor supplied by a Python host.

Other interfaces

CLI

Use the CLI for manual operation, automation, artifact inspection, or when an agent host can read and write the stage YAML files but cannot call MCP directly. The guided workflow is mirrored as llm-migrate run start|status|tasks|blockers|decide|submit-prompt|submit-file| confirm-unaffected|record-validation|finalize|review|decide-change, plus llm-migrate models match and llm-migrate research prompts.

llm-migrate --help
llm-migrate run --help
llm-migrate research --help
llm-migrate plan --help
llm-migrate eval --help

The CLI and MCP server are thin interfaces over the same Python service.

Headless Python

A Python host can implement llm_migrate.core.orchestration.AgentRunner and call MigrationService.run_agent_research(...). The host supplies agent calls; the service still owns budgets, retries, resumability, artifacts, and gates.

Agent workflow instructions

docs/agent-research-workflow.md is the agent-neutral workflow used by Codex, Claude Code, or another host. It describes how the host should run research and review. It is not a standalone model or hosted service, and it is not packaged as a separate installed Codex or Claude skill. The MCP server is the primary agent-facing product interface.

Safety and trust

  • Research agents receive exact identities and normalized requirements rather than the full repository whenever possible.

  • Repository and webpage content is untrusted data and cannot change orchestration instructions, permissions, limits, or schemas.

  • A researcher cannot review its own claims.

  • Contradictory evidence remains unresolved; majority vote does not make it true.

  • Session knowledge is hash-linked, expiring, user-scoped, and visibly session_agent_reviewed or session_unreviewed.

  • Only a maintainer-reviewed proposal can change the canonical registry.

  • Planning, preparation, and adaptation never rewrite application source files; deliverables live in the run workspace.

  • The agent presents blocker questions, options, evidence, and review changes verbatim; the user decides, and an explicit accept requires the user's own rationale.

  • Every edit in a deliverable is an anchored, evidence-linked annotated change reconciled against the real diff; serialization tricks cannot hide content from validation or make an unchanged deliverable look adapted.

  • The generated contract test is emitted for you to run; the toolkit never executes your application, tests, or provider calls on its own.

research -> evidence -> independent review -> session overlay -> migration plan
                                            \
                                             -> maintainer proposal -> canonical registry

Current scope and limitations

Area

Current behavior

Application scanning

Python-focused, plus structural YAML/JSON/TOML prompt-configuration discovery

Built-in registry

Intentionally small and evidence-backed; synthetic profiles are labeled TEST/FIXTURE

Prompt generation

Deterministic candidate preparation in core; model-authored rewrites belong to the agent host and face dispositions, annotations, and review

Source mutation

No automatic code or prompt rewriting

Dynamic inputs

Unresolved dynamic prompts or configuration remain explicit unknowns

Research agents

Supplied and paid for by the user's host

Agent orchestration

Sequential today; persistent caching and orchestrated arbitration are deferred

Network access

Explicit research/search in the host, live pricing, runtime evaluation, and cited-source refetching

Infrastructure

No hosted backend, telemetry, database, or project-owned credentials

Contributing

Contributions are welcome through reviewed pull requests. Start with CONTRIBUTING.md, follow the evidence requirements in the research policy for registry work, and report security issues privately through SECURITY.md.

Project governance is documented in GOVERNANCE.md, and all participants must follow the Code of Conduct.

Development

python -m pip install -e '.[dev]'
pytest
ruff check .
ruff format --check .
mypy

The repository registry is discovered automatically. To use another registry, pass --registry PATH or set LLM_MIGRATE_REGISTRY to a root containing models/.

Keeping the registry fresh

Freshness windows are short by design (7 days for pricing and lifecycle, 14 for availability, 30 for capabilities and prompt guidance). Maintainers keep them honest with a refetch-only refresh that never edits registry/:

llm-migrate registry refresh-evidence --model claude-sonnet-5

It refetches only the source URLs already recorded in the profile, hashes each page's visible text against the baseline in .registry-proposals/<model>/refresh-evidence.yaml, and reports: unchanged pages that still name the model propose moving checked_at for the categories they support; a page that no longer names the model (a lineup page that dropped a legacy model) vouches for nothing; changed pages are held until you re-verify them (then --rebaseline); a URL with no baseline gets one recorded and proposes nothing; a category the bundle's review put on hold stays held. Promote an approved proposal by hand through the model's bundle, like every other registry change. --no-record produces a read-only report; --output DIR writes it as YAML and Markdown.

Design references:

License

Copyright © 2026 Athena Li.

Licensed under the Apache License, Version 2.0. See LICENSE.

Available Tools

47 tools
add_prompt_sourcesAdd Prompt SourcesA

Add prompt source files to a LIVE run, or dismiss discovery candidates.

paths are application-relative prompt files the user confirmed are live prompts; the worklist re-derives (submitted deliverables keep their entries) and new_prompt_tasks lists the added pending tasks. dismiss names unreferenced candidate files or dynamic consumer path:line:keyword addresses the USER says are not prompts / genuinely runtime-built, with the user's rationale (required). Nothing is written on any problem.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
pathsYes
dismissNo
run_dirYes
rationaleNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and delivers meaningful disclosure: the operation targets a 'LIVE' run (indicating mutation), the worklist 're-derives (submitted deliverables keep their entries),' and critically, 'Nothing is written on any problem' — an atomic/no-op-on-failure guarantee. It omits permissions and reversibility, but covers the most decision-relevant effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized: a one-sentence purpose statement up front, followed by tight backtick-delimited parameter semantics and a closing atomicity guarantee. Every sentence adds information; nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists and there are no annotations, the description covers the key ground: what each mode does, what effects occur (re-derivation, new_prompt_tasks), and failure behavior. Remaining gaps are minor — the unexplained `now` parameter and no explicit sibling routing — but the tool is a moderately complex 5-parameter mutation and is described well enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does for the substantive parameters: paths (app-relative, user-confirmed), dismiss (candidate files or path:line:keyword addresses), and rationale (required for dismiss). The gap is `now`, which is never mentioned, and `run_dir` is only implied via 'application-relative' rather than explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names specific verbs and resources — 'Add prompt source files to a LIVE run, or dismiss discovery candidates' — and usefully reveals the dismiss mode that the title alone hides. This distinguishes the tool from the many submission/confirmation siblings (submit_adapted_prompt, confirm_prompt_consumer, get_research_prompts) by its live-run worklist context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit conditions for each operational mode: paths are 'application-relative prompt files the user confirmed are live prompts,' while dismiss targets 'unreferenced candidate files or dynamic consumer addresses the USER says are not prompts' with a required rationale. It does not name alternatives or exclusions among siblings, but the within-tool usage context is clear and unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_invocationAnalyze InvocationD

Analyze provider invocation and adjacent tool/output couplings.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNo
target_platformNo
application_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavior. It only says 'analyze', which implies read-only but does not state whether it has side effects, requires special permissions, or what side effects might occur. No behavioral detail beyond the verb is offered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, but it is under-specified to the point of uselessness. This is not effective conciseness; it is a lack of necessary content. The single sentence does not front-load any actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description is completely inadequate for a tool with 3 parameters and no annotations. It does not explain the tool's purpose, when to use it, parameter meaning, or any behavioral context, leaving the agent with no basis for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the three parameters (target, target_platform, application_path). The mention of 'provider invocation' and 'couplings' does not map to any parameter semantics, so the agent has no clue how to fill them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb 'analyze' and a resource 'provider invocation and adjacent tool/output couplings', but this is vague and does not clarify what specific analysis is performed. It does not distinguish from sibling analysis tools like analyze_prompt or analyze_regressions, and the term 'couplings' is undefined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus any of the 32 siblings. There is no mention of conditions, prerequisites, or alternatives, leaving the agent to guess based solely on the name and vague description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_promptAnalyze PromptC

Conservatively analyze prompt characteristics without inference.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It conveys a conservative, no-inference approach, but says nothing about side effects, return behavior, limitations, or permissions. This is only a thin behavioral hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no redundancy, and the core idea is front-loaded. However, it is under-specified rather than efficiently complete, so it earns only a midscore.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not strictly required. Still, the description gives an agent no way to choose this tool among many prompt-related siblings given the large sibling list, and the absence of annotations compounds the issue.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no parameter-level detail. An agent only sees a string named 'prompt' with no explanation of expected format, content, or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Describes a specific action (analyze) on a specific resource (prompt characteristics), and adds a distinctive qualifier ('conservatively', 'without inference'). It doesn't explicitly differentiate from sibling tools like validate_prompt or analyze_invocation, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance about when to use this tool versus alternatives. The description does not state what scenarios warrant conservative analysis, nor does it name any sibling tools to prefer for validation, comparison, or inference-based tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_regressionsAnalyze RegressionsB

Return a structured categorized regression report for an evaluation run.

ParametersJSON Schema
NameRequiredDescriptionDefault
runYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. 'Return a structured categorized regression report' signals a read-only reporting behavior, but it does not disclose side effects, prerequisite data requirements, or how categorization is determined. This is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one clean, front-loaded sentence with no wasted words. It immediately states the action and the output, which is ideal for the tool's simple interface.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and may document the report structure, the input side is underspecified. The single required 'run' parameter is an unconstrained object, and the description does not clarify what an evaluation run must contain or how to obtain it. This leaves meaningful ambiguity for an agent trying to invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the schema only provides a generic 'run' object with additionalProperties true. The description adds that 'run' is an evaluation run, which is useful, but gives no detail about required fields, expected structure, or how the run object should be supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Return'), resource ('structured categorized regression report'), and target ('evaluation run'). It is clear and focused, though it does not explicitly contrast with similar reporting/analysis siblings such as analyze_prompt or generate_migration_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear context for use: analyzing regressions from an evaluation run. It does not explicitly state when not to use it or name alternative tools, but the intended scenario is reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_research_consensusBuild Research ConsensusC

Deterministically combine research and review; agent agreement is not evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
reviewYes
requestYes
researchYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, yet it only reveals that the combination is deterministic and that agent agreement is not treated as evidence. It does not disclose side effects, required permissions, return behavior, or what 'combine' produces beyond the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise and front-loaded with the core operation, and the warning is packed into a short second clause. It is not bloated, though the brevity leaves important information unwritten.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite an output schema existing, the description is incomplete for a tool with three required nested-object parameters and zero parameter documentation. An agent would know the general intent but not how to structure the inputs or what behavioral guarantees apply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the three parameters have descriptions. The text references 'research' and 'review' but never clarifies their structure or roles, and 'request' is entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('deterministically combine') and the resources ('research' and 'review'), and the title 'Build Research Consensus' supplies the intended result. It is not a tautology and is distinguishable from siblings like validate_evidence_review, though the output format is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, no preconditions, and no named sibling. The caveat 'agent agreement is not evidence' hints at a use context but does not tell an agent when to choose this tool over validate_research_result or generate_migration_research_request.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_session_registryBuild Session RegistryA

Finalize a run workspace into an immutable, expiring session overlay.

Research and review artifacts must already exist in the workspace; any stage still needing an agent fails closed instead of doing hidden work.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes
ttl_daysNo
shadow_canonicalNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must cover behavioral traits. It discloses that the tool fails closed if stages still need an agent, and implies it does not do hidden work. It mentions immutability and expiring overlay, which are significant behavioral traits. However, it does not detail what failure returns or specific side effects beyond the fail-closed behavior, but the core transparency is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: the first line states the core purpose and key characteristics (immutable, expiring overlay). The second line provides crucial prerequisites and failure behavior. It is well-structured with no redundant information, every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (finalizing a session overlay with guardrails) and the existence of an output schema, the description is fairly complete. It covers the main requirements (artifacts exist, fail-closed) and outcomes (immutable, expiring). It lacks details on parameter semantics and specific return behavior, but the presence of an output schema reduces the need for return documentation, and the essential context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no parameter-specific details. The tool has 4 parameters (run_dir, now, ttl_days, shadow_canonical) with a required run_dir, but the description doesn't explain the meaning of any beyond implicit inference from the description (e.g., ttl_days likely relates to expiration). This is a gap that forces the agent to guess at parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to finalize a run workspace into an immutable, expiring session overlay. It identifies the specific verb ('finalize') and the resource ('run workspace' to 'session overlay'), which helps distinguish it from other migration-related tools. However, it does not explicitly differentiate itself from sibling tools like 'finalize_migration', requiring the agent to infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when to use: artifacts must already exist, and any stage needing an agent fails closed. It implies this is for the final step after research/review artifacts are prepared. It does not explicitly mention when not to use it or list alternatives, but the prerequisites are clearly stated, making the usage context clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_model_lifecycleCheck Model LifecycleB

Interpret reviewed lifecycle facts at a fixed date without live research.

ParametersJSON Schema
NameRequiredDescriptionDefault
as_of_dateNo
identifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the tool does not perform live research and works from reviewed facts, which implies a read-only, offline behavior. However, it does not mention side effects, staleness, or failure behavior, leaving some transparency gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused sentence with no redundancy or filler. It is appropriately compact for a simple tool, though the terseness leaves some semantic gaps.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the tool lacks annotations and has zero parameter documentation in the schema. The description does not define 'lifecycle facts,' explain the identifier input, or give enough selection criteria among the many sibling tools. It is not complete enough for an agent to confidently invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. It only hints at 'a fixed date' (matching as_of_date) but never explains the required 'identifier' or the format/meaning of the date. This is insufficient for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Interpret') and resource ('reviewed lifecycle facts'), and adds a fixed-date scope plus 'without live research.' It is clear enough to distinguish the tool from live-research siblings, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without live research' provides clear contextual guidance: use this tool when you need reviewed lifecycle facts at a fixed date rather than live research. It does not explicitly list alternatives or when-not-to-use conditions, but the context is reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_modelsCompare ModelsC

Return explicit same/different/unsupported/unknown comparison states.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
targetYes
source_endpointNo
source_platformNo
target_endpointNo
target_platformNo
include_live_pricingNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does reveal the output contract (explicit discrete states, including unsupported and unknown), which is useful, but it says nothing about whether the tool has side effects, calls external services such as live pricing, requires auth, or how unsupported/unknown are determined.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single-sentence description is tightly written, front-loads the key result (explicit comparison states), and has no filler. It is concise but under-specified; the short length is appropriate, yet the one sentence cannot carry the full guidance burden.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the tool still lacks essential context: 7 parameters with 0% schema documentation, no annotations, and no description of how source/target/endpoint/platform/pricing combine. An agent cannot confidently construct a correct call from the available text alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description needed to explain what source/target mean and how endpoint/platform/pricing parameters affect the comparison; it does none of that. The property names are somewhat self-explanatory, but the text adds no semantic value beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and names the explicit output states (same/different/unsupported/unknown), so an agent can tell this is a discrete-state comparison tool rather than a free-form diff. It does not literally state that source and target are models, but the title and sibling context make that reasonably clear. It stops short of 5 because compare_outputs could be confused without more explicit object-of-comparison wording.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to choose this tool over compare_outputs, resolve_model, or the other comparison/migration siblings, and no mention of prerequisites or exclusions. The description only states what the tool returns, leaving the agent to infer when it is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_outputsCompare OutputsC

Compare one paired source/target evaluation result deterministically.

ParametersJSON Schema
NameRequiredDescriptionDefault
source_resultYes
target_resultYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'deterministically' which is a behavioral trait, but does not disclose whether the operation is read-only, has side effects, requires specific permissions, or what the result structure is. This is a significant gap for a tool with no annotation safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence, which is concise and direct. However, it is under-specified, lacking necessary details about inputs and behavior. It is not verbose, but the brevity comes at the cost of completeness, making it borderline adequate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema means return values are covered, but the description omits critical context: when to use this tool, what the parameters mean, and any behavioral notes. Without annotations or parameter descriptions, an agent would struggle to call this correctly. It is incomplete for a two-parameter tool with no safety annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, meaning the parameters are undocumented in the schema. The description only hints that the two parameters form a 'paired source/target' relationship, which provides minimal meaning. It does not explain what constitutes a source or target result or how they should be structured. The description fails to compensate for the lack of parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('compare') and a specific resource ('one paired source/target evaluation result'). This is clear enough to distinguish it from siblings like compare_models, which likely compare models rather than evaluation results. However, it does not elaborate on what 'compare' entails or what the output looks like, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention conditions, exclusions, or provide any context for selection. An agent would have to infer that it is for comparing a pair of evaluation results, but no explicit usage guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_prompt_consumerConfirm Prompt ConsumerA

Record that one dynamic prompt consumer reads one prompt file.

location is a path:line:keyword from the run's dynamic_prompt_consumers (get_run_status / list_adaptation_tasks); source_path is the file the USER confirmed it reads. The consumer becomes source-backed, the file a prompt source, and the worklist re-derives.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes
locationYes
source_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full disclosure burden, and it performs well: it states the concrete state transitions ('The consumer becomes source-backed, the file a prompt source, and the worklist re-derives'), which tells the agent this is a mutating operation with side effects on the run's worklist. It does not cover idempotency, overwrite behavior, or error conditions, but the core behavioral contract is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short paragraphs with zero filler. The purpose sentence is front-loaded, and the second paragraph packs parameter semantics and state-change consequences into three tightly-worded sentences. Every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. The description covers purpose, workflow context, parameter semantics, and side effects, which is enough for an agent to call it correctly after querying get_run_status/list_adaptation_tasks. Minor gap: the optional `now` parameter's role is never explained, though its null default softens the impact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for the two domain-critical parameters: `location` is specified as a `path:line:keyword` with provenance from dynamic_prompt_consumers, and `source_path` is defined as 'the file the USER confirmed it reads.' `run_dir` is self-evident from its name, though `now` remains unexplained; the two params requiring domain knowledge are well covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Record that one dynamic prompt consumer reads one prompt file.' It then states the precise outcome ('consumer becomes source-backed, the file a prompt source, and the worklist re-derives'), which clearly separates it from siblings like add_prompt_sources or confirm_unaffected. The purpose is unambiguous and action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear workflow context: `location` must come from the run's dynamic_prompt_consumers as reported by get_run_status / list_adaptation_tasks, and `source_path` is a file the USER confirmed. This tells an agent exactly when in the migration workflow this tool fits. It stops short of a 5 because it never names alternatives explicitly or states when NOT to use it versus add_prompt_sources.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_unaffectedConfirm UnaffectedA

Close every listed unaffected_files entry with one reviewed no-change call.

Each path is validated independently (per-file accept/reject in results): only files on the worklist's unaffected_files list qualify — a file with required changes needs a real submission. Each accepted path records a reviewed no-change entry that appears in the final report. acknowledge_source_references=true also closes files finalize flagged source_reference_uncovered whose source-model mention the USER says is intentional (docs, history); pass the user's own rationale.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
pathsYes
run_dirYes
rationaleYes
acknowledge_source_referencesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses per-file independent validation, accept/reject behavior in `results`, the condition under which `acknowledge_source_references` closes `source_reference_uncovered` files, and that accepted entries appear in the final report. It could additionally state what happens when a path is rejected, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The primary purpose is front-loaded in the first sentence, and the subsequent sentences add genuinely useful qualification and behavioral detail. The density is high and each sentence earns its place, though the phrasing is somewhat jargon-heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists and there are no annotations, the description explains the tool's qualification rules, validation behavior, report impact, and the optional acknowledgment mode. The main omissions are the semantics of `run_dir` and `now`, but the essential calling context is well covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaningful semantics for `paths` (only unaffected entries qualify), `acknowledge_source_references` (closes source_reference_uncovered files when the user confirms intent), and `rationale` (pass the user's own rationale). However, `run_dir` and `now` are not explained at all, leaving gaps for two of the five parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource statement: 'Close every listed `unaffected_files` entry with one reviewed no-change call.' It clearly distinguishes this action from related sibling tools by defining the exact scope (unaffected files only) and the no-change nature of the call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit inclusion/exclusion guidance: only files on the worklist's `unaffected_files` list qualify, and a file with required changes needs a real submission rather than this no-change call. It also explains when `acknowledge_source_references=true` is appropriate, though it does not name a specific alternative sibling tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_migration_research_requestCreate Migration Research RequestC

Scan locally and bound explicit user-scoped research to missing/stale facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofYes
run_idYes
sourceYes
targetYes
topicsNo
applicationYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. It does not state whether this creates a persistent record, what side effects occur, what permissions are needed, or what the output represents. 'Scan locally' suggests a read-like action, while the tool name implies creation, leaving behavior unclear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but under-specified. It does not front-load the core purpose or add usable information; the single sentence is more cryptic than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, nested objects, an output schema, no annotations, and many sibling tools, this description is far too incomplete. It omits what a migration research request is, how the parameters relate, when to use it, and what the tool actually accomplishes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not clarify any of the six parameters: application, run_id, source, target, as_of, or topics. Phrases like 'locally' and 'user-scoped' are not mapped to the schema, so the agent gets no meaningful parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description contains action verbs ('Scan', 'bound') and hints at an operation on research facts, but it never explicitly says it creates a migration research request or what that entails. It is vague rather than tautological, and it does not distinguish this tool from any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool instead of the many sibling tools. 'Scan locally' implies some local scope, but no explicit conditions, prerequisites, or alternatives are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

estimate_migration_costEstimate Migration CostA

Estimate recurring token cost from checked-in canonical pricing.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
targetYes
requestsYes
input_tokens_per_requestYes
output_tokens_per_requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral burden. It does disclose that the estimate is recurring and uses canonical/checked-in pricing, implying read-only calculation behavior. However, it does not clarify side effects, assumptions, or what happens if pricing data is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every word adds meaning: 'recurring,' 'token cost,' and 'checked-in canonical pricing' all convey scope and data source.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five required parameters, no annotations, and no schema descriptions, the one-sentence description is not sufficient for an agent to invoke the tool confidently. The output schema exists but does not compensate for the missing input semantics and behavioral context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the five parameters. Parameter names like source, target, requests, and token counts are suggestive, but the description adds no units, constraints, or clarification about what values are expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation: 'Estimate recurring token cost' and scopes it to 'checked-in canonical pricing.' This clearly distinguishes the tool from live-pricing siblings like query_live_pricing, and it names the resource being estimated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'From checked-in canonical pricing' gives clear context for when to use this tool: when the estimate should be based on the static/checked-in pricing source rather than live quotes. It does not explicitly name alternatives or exclusions, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

finalize_migrationFinalize MigrationA

Write the run's manifest, report, and adaptation change log under output/.

The report records per-file changes and rationale, states reviewed no-change deliverables explicitly, lists remaining coverage gaps, and runs the cross-surface consistency gate over the whole deliverable set. Re-run it any time; it always reflects the current submissions.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure. It discloses idempotency ('Re-run it any time') and the consistency gate, but it does not describe whether files are overwritten, what happens on gate failure, or any side effects beyond writing output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all load-bearing: main action, report contents, and idempotency. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations and an output schema, so the description covers core behavior (writes artifacts, consistency gate, idempotency). Missing are prerequisites (e.g., that submissions should already be recorded), failure behavior if the gate fails, and how output files relate to the run directory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain run_dir or now. The term 'run's' implies run_dir, but no format, meaning, or relationship to current submissions is provided, so it fails to compensate for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: writes the run's manifest, report, and adaptation change log under output/. It also details the report's contents and mentions the consistency gate, which distinguishes it from siblings like generate_migration_report or start_migration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: the tool finalizes a run's artifacts and can be re-run at any time to reflect current submissions. It does not explicitly name sibling alternatives or state when not to use it, so not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_eval_suiteGenerate Eval SuiteC

Bind a deterministic evaluation corpus to a migration manifest.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNomigration-evaluation
casesYes
migration_planYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of disclosing side effects and behavior. It only characterizes the corpus as 'deterministic,' but does not state whether this mutates state, writes files, validates inputs, or is idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is efficient and contains no filler, which is good for front-loading. However, it is under-specified to the point that conciseness comes at the expense of usefulness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists, the tool has nested, loosely constrained objects (cases and migration_plan are additionalProperties objects), no annotations, and no usage context. The description is far too thin for an agent to invoke the tool correctly or understand what the resulting evaluation suite contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description was expected to compensate, but it only alludes to 'evaluation corpus' for cases and 'migration manifest' for migration_plan. It adds no detail about the contents of cases, the structure of migration_plan, the optional name parameter, or how parameters interrelate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Bind') and two concrete resources ('deterministic evaluation corpus', 'migration manifest'), which makes the core operation clear and distinguishes it from siblings like run_migration_eval or generate_migration_plan. However, 'bind' is somewhat opaque jargon and alternatives are not named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus related tools such as run_migration_eval, compare_outputs, or generate_migration_plan. There are no preconditions, no exclusions, and no workflow context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_migration_planGenerate Migration PlanB

Generate an actionable application-level migration manifest without writing files.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
targetYes
prompt_sourcesNo
source_endpointNo
source_platformNo
target_endpointNo
target_platformNo
application_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does state that the tool does not write files, which is a key side-effect trait, but it does not mention whether it reads files, requires network access, or any other behavioral aspects. The description is truthful but minimal, earning a middle score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no fluff, and the key information (application-level, no file writing) is front-loaded. While it could be longer to address other dimensions, the conciseness itself is appropriate; the issue is more about completeness than structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 8 parameters, no schema descriptions, and no annotations, the description is far too sparse. It does not explain the meaning of 'actionable,' the nature of the manifest, or how parameters influence the result. Although an output schema exists (so return values need not be described), the lack of usage and parameter guidance makes the definition incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage and the tool description does not explain any of the 8 parameters (e.g., application_path, source, target, or optional fields). Since the description must compensate for the schema's lack of information, this complete omission is a severe deficiency.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates an 'actionable application-level migration manifest' and explicitly notes it does so 'without writing files.' This specifies the verb, resource, and scope, and differentiates it from siblings like generate_migration_report (report) and generate_session_migration_plan (session-level) by the 'application-level' qualifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. There is no mention of conditions, exclusions, or references to sibling tools, leaving the agent to infer usage solely from the name and generic description. With many closely related tools in the sibling list, this is a notable gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_migration_reportGenerate Migration ReportC

Generate a human-readable report from the integrated migration workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
targetYes
prompt_sourcesNo
source_endpointNo
source_platformNo
target_endpointNo
target_platformNo
application_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose side effects and operational traits. It says 'generate a human-readable report', which implies a read-like operation, but it never states whether this tool actually runs migrations or modifies anything. It also fails to mention any safety, auth, or side-effect details. The disclosure is minimal and incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no fluff or redundancy. However, it is under-specified and does not provide enough structural detail to be genuinely useful. It is appropriately sized in length but not in content, so it only partially earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a tool with 8 parameters (3 required) and no output schema visible, the description is drastically incomplete. It does not explain what the 'integrated migration workflow' entails, what inputs are needed, what the report contains, or any relationship to the migration process. An agent has virtually no context to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description adds no meaning for any of the 8 parameters. It does not explain what source, target, application_path, or the optional fields represent, nor how they are used to generate the report. The burden falls entirely on the description, and it fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a verb (generate) and a resource (human-readable report), and it conveys the output type. However, it does not differentiate from sibling tools like generate_migration_plan or generate_session_migration_plan, which also produce migration-related outputs. It is clear but lacks distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the many sibling migration tools. It does not mention any preferred scenario, exclusions, or alternatives. An agent receives no help in selecting between this and other plan/report tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_session_migration_planGenerate Session Migration PlanC

Plan a migration over the canonical registry plus one run's session overlay.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
sourceYes
targetYes
applicationYes
session_run_dirYes
source_platformNo
target_platformNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must convey side effects and behavior. 'Plan' hints at non-execution, but it never states whether any mutation occurs, what inputs are needed, how the overlay affects the result, or what the plan contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with no filler, and the core purpose is front-loaded. However, conciseness comes at the cost of omitting useful behavioral and parameter context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with seven parameters, four required, and no annotations, this description is too thin. It provides no guidance on how to invoke it, what makes this session-based version different from alternatives, or what the output plan covers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description offers no explanation of the parameters. Terms like 'canonical registry' and 'session overlay' do not map to source, target, application, or session_run_dir, leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific action and resource: planning a migration over the canonical registry plus one run's session overlay. It is not the same as a generic migration plan because of 'session overlay,' which helps distinguish it slightly from siblings like generate_migration_plan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the 'session overlay' phrasing suggests this is for migration planning involving a single session run, unlike the general generate_migration_plan. There is no explicit when-to-use/when-not-to-use guidance or named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_blocker_resolutionsGet Blocker ResolutionsA

Questions plus evidence-backed options for every unresolved blocker.

Each blocker carries the question to ask the user and registry-backed options with consequences, evidence URLs, and the exact record_blocker_decision call. Present them VERBATIM, one blocker at a time; never choose on the user's behalf.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond a simple 'get' by detailing the presentation behavior: 'Present them VERBATIM, one blocker at a time; never choose on the user's behalf.' It also discloses the content of the options (consequences, evidence URLs, and the exact record_blocker_decision call), which gives the agent a clear understanding of what to expect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with zero fluff. The first sentence front-loads the core purpose, and the subsequent sentences add necessary detail about content and presentation rules. Every sentence earns its place, and the structure is clear and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the output schema exists, its content is not provided, so the description must carry the burden. It explains what the output contains (questions, options, evidence URLs, call) but completely omits parameter semantics, which is critical for an agent to invoke the tool correctly. The description is also silent on error cases or prerequisites, leaving the agent with incomplete guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate by explaining the parameters. It fails to mention 'now' or 'run_dir' at all, leaving their purpose and expected values entirely undocumented. An agent cannot infer what to pass for either parameter based on the description, making this a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: it returns questions and evidence-backed options for unresolved blockers. It specifies the resource (blockers) and the action (get), and it distinguishes itself from the sibling record_blocker_decision by emphasizing that it presents options verbatim and never chooses on the user's behalf, which is the opposite of recording a decision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when this tool should be used (when there are unresolved blockers to resolve) but does not explicitly state when to use it versus alternatives. It mentions the exact record_blocker_decision call but does not explicitly say 'use this tool to retrieve options, then use record_blocker_decision to record the choice,' leaving the workflow inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_change_reviewGet Change ReviewA

Per-deliverable annotated changes paired with their decision state.

Returns each deliverable's decoded diff and every annotated change (why, evidence, spans, pending/accepted/rejected). Present each pending change VERBATIM, one at a time; the user decides, never you. Drift and stale decisions are reported, never silently applied.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden, and it delivers: it states that pending changes must be shown verbatim, the agent must never decide, and drift/stale decisions are reported rather than silently applied. This gives an agent a clear mental model of the tool's side-effect-free, user-decision-oriented behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three focused sentences: a summary, a return-content enumeration, and a behavioral instruction. It is front-loaded, avoids repetition, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and the description richly covers return semantics and decision handling, it leaves the only input parameter completely undocumented and provides no explicit usage context beyond the implied 'present pending changes.' For a tool with a single required parameter, this is a significant completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole required parameter, run_dir, has zero schema description coverage and is never explained in the description. The description does not compensate for this gap by stating what run_dir means, what format it should use, or how to obtain it, leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it returns per-deliverable annotated changes paired with decision state, including decoded diff and every annotated change. This clearly distinguishes it from siblings like record_change_decision and get_run_status, which involve writing decisions or reporting status rather than retrieving review content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The instruction to present each pending change verbatim and let the user decide implies the tool is used when delivering pending changes for human judgment. However, it does not explicitly name alternative tools or state when not to use it, so the usage guidance is implied rather than fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_model_profileGet Model ProfileC

Return a validated local registry model profile.

ParametersJSON Schema
NameRequiredDescriptionDefault
identifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states that it returns a profile, but does not explain side effects (likely none), performance characteristics, whether it errors if the identifier is not found, or the semantics of 'validated' and 'local registry'. This is insufficient for a tool with no annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (one sentence), but conciseness is not the same as adequacy. It is not front-loaded; it is just a minimal phrase that could be a tautology. It is not padded with fluff, but it is under-specified, which is a different issue. A score of 3 is generous because it is at least short.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the tool has only one parameter and an output schema, the description fails to clarify the tool's role in the larger migration workflow. Given the sibling names, this tool likely serves a specific function in model profiling, but the description does not explain what 'validated' or 'local registry' mean, nor how it relates to other tools. The output schema might help, but it is not visible to the agent without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema coverage is 0% and the description does not explain the 'identifier' parameter beyond the schema's type. It does not mention that the identifier likely refers to a model name or ID, or how it should be formatted. With a single parameter, the description should at least clarify what the parameter represents, but it adds no value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is minimal: 'Return a validated local registry model profile.' It specifies a verb ('Return') and a vague resource ('model profile'), but it does not define what 'validated' or 'local registry' means, nor does it distinguish from siblings like 'resolve_model' or 'compare_models'. It merely restates the tool name without useful detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives such as 'resolve_model' or 'check_model_lifecycle'. It does not mention any context, prerequisites, or exclusions. An agent would have no idea whether to call this or a sibling for a given task.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_research_promptsGet Research PromptsA

Ready-to-run researcher and reviewer prompts for a run's research request.

One bounded prompt pair per remaining scope. Run each with a separate agent (never one agent for two scopes or reviewing its own research), write the YAML artifacts to the stated paths, validate each with validate_research_artifact, then call build_session_registry.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description discloses the important behavioral constraints: each prompt pair is for a separate scope, must run under separate agents, and must be validated before registry build. It does not explicitly state whether the call itself has side effects, but 'Ready-to-run prompts' implies retrieval, and the output schema covers return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences front-load the purpose and then provide the necessary workflow constraints. Every clause earns its place; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, scoping, agent constraints, artifact handling, validation, and the next registry step, which is substantial for a single-parameter tool. It leaves 'remaining scope' and 'stated paths' slightly underspecified, but the output schema likely supplies those details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, run_dir, with 0% schema description coverage. The description says the prompts are 'for a run's research request,' which implies run_dir identifies the run, but it never explicitly defines run_dir as the path or identifier, leaving the agent to infer it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete deliverable ('researcher and reviewer prompts for a run's research request') and bounds it with 'one bounded prompt pair per remaining scope,' making the tool's function clear. It also distinguishes the tool from sibling research/validation tools by naming the follow-up tools (validate_research_artifact, build_session_registry) rather than confusing it with generation or analysis tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit operational context: run each prompt pair in a separate agent, never have an agent review its own research, write YAML artifacts, validate, then call build_session_registry. It does not explicitly state when not to use this tool or name alternatives, but the research-request context and workflow are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_run_statusGet Run StatusA

The run's state machine position with the single next action.

States: research_pending -> discovery_incomplete -> blockers_pending -> tasks_pending -> review_pending -> ready_to_finalize -> validation_pending, with pending counts. discovery_incomplete (prompt coverage unresolved with candidate files, or any unresolved consumer in strict mode) names the exact add_prompt_sources / confirm_prompt_consumer calls. Served from the worklist snapshot (re-derived only when the application, run identity, decisions, or registry changed), so it is cheap to call between steps.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well. It discloses that results come from a worklist snapshot, lists the conditions that force re-derivation, states that calls are cheap, and explains the special behavior for the discovery_incomplete state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core meaning, followed by a compact state list, a necessary special-case explanation, and a caching note. Every sentence adds distinct value and the structure makes the state machine easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers the return shape, and the description thoroughly explains state semantics and snapshot behavior. The only real gap is the lack of any input-parameter guidance, especially for the optional now parameter, so the description is strong but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention either parameter. The meaning of run_dir and especially the optional now parameter is left entirely to the agent to infer, so the description provides no compensating guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as returning a run's state machine position plus the single next action, and it enumerates the exact state sequence. This goes well beyond the bare title and makes the tool's role unambiguous relative to its many siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear contextual guidance that the tool is served from a snapshot and is cheap to call between steps, which tells the agent when it is appropriate to invoke it. It does not explicitly state when not to use it or name a direct alternative, but the usage context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_adaptation_tasksList Adaptation TasksA

Per-file adaptation worklist for a started migration run. Call it ONCE.

shared_prompt_guidance applies to every prompt task; guidance items carry stable ids for guidance_dispositions. Each prompt task's verbatim_source is the ORIGINAL unmodified content, never a proposed adaptation. Do not re-list between submissions; finalize_migration reports remaining gaps.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It adds valuable context about the data returned: shared_prompt_guidance applies to all prompt tasks, guidance items have stable ids, and verbatim_source is always original content. It also cautions against re-listing. It does not explicitly state side-effect-freedom, but the list semantics strongly imply a read-only operation, so the transparency is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficient and front-loaded with the purpose and the critical 'Call it ONCE' instruction. It uses a compact paragraph with clear sub-points, every sentence adds value, and there is no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description need not explain return format. It covers key data semantics (stable ids, original source) and usage constraints. The only significant gap is parameter explanation, but the tool is simple with just two parameters, and the description provides enough context for an agent to understand run_dir's role. Overall, it is sufficiently complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description does not explain either parameter. It implies run_dir identifies the migration run but never defines it explicitly, and the 'now' parameter is entirely unmentioned. The description should compensate for the lack of schema documentation but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose: 'Per-file adaptation worklist for a started migration run.' It names the resource (adaptation tasks) and the context (started migration run), making it distinct from sibling tools like start_migration or finalize_migration. However, it does not explicitly name alternative tools, so it falls short of the strongest differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Call it ONCE.' and 'Do not re-list between submissions; finalize_migration reports remaining gaps.' This tells the agent exactly when to use the tool and when to switch to a sibling, leaving no ambiguity about invocation frequency or fallback behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimize_migrationOptimize MigrationB

Return bounded, review-only recommendations from observed regression evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitsNo
candidate_runsNo
regression_reportYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It explicitly says 'review-only' and 'Return ... recommendations,' which signals a non-mutating, analysis-oriented tool. However, it does not disclose output limits' behavior, potential errors, or any side effects beyond that, leaving moderate gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It front-loads the core behavior ('Return bounded, review-only recommendations') and then adds the evidence source, making it appropriately sized and efficiently structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given three parameters, zero parameter descriptions, no annotations, and a large sibling set, this description is insufficient for correct invocation. It does not explain the optional parameters, nor does it clarify how this tool differs from related migration-planning or analysis tools, despite having an output schema to cover return shapes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the opaque parameters. It only hints at 'regression evidence,' loosely mapping to the required regression_report, but gives no meaningful explanation of limits or candidate_runs. The agent cannot infer what shapes or semantics those optional parameters expect.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Return') and resource ('bounded, review-only recommendations') tied to 'observed regression evidence,' so an agent can infer the tool produces evidence-based suggestions. However, it does not explicitly differentiate from sibling tools like analyze_regressions or generate_migration_plan, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'from observed regression evidence' implies the tool should be used when a regression report is available and review-only recommendations are needed. There is no explicit when-to-use/when-not-to-use guidance or mention of alternatives among the many migration-related siblings, so the guidance remains implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_invocation_migrationPrepare Invocation MigrationC

Prepare a target invocation contract without editing application files.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
targetYes
source_platformNo
target_platformNo
application_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It only says 'prepare' and 'without editing application files,' which is vague—it doesn't explain what the tool actually does (e.g., creates a contract object, returns data, requires permissions) or what side effects might occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with a clear verb and subject. It is front-loaded with the primary action and includes a relevant constraint, making it easy to parse. However, its brevity comes at the cost of completeness, which is penalized elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description is grossly incomplete for a tool with five parameters and no annotations. It lacks parameter explanations, usage context, and any indication of when to call it in the migration workflow. An agent would not know what inputs to provide or how this tool fits with siblings like prepare_prompt_migration or generate_migration_plan.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no information about the five parameters (application_path, source, target, source_platform, target_platform). The names provide some hints, but the description does nothing to clarify their meaning, types, or relationships, forcing the agent to infer from names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('Prepare a target invocation contract') and adds a key differentiator ('without editing application files'), which distinguishes it from file-editing tools. However, it doesn't explicitly distinguish from similar siblings like prepare_prompt_migration, so it's not a perfect 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. No mention of prerequisites, conditions, or exclusions. The description gives no context about the migration workflow, leaving the agent to guess whether this should precede generate_migration_plan or start_migration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_prompt_migrationPrepare Prompt MigrationB

Prepare a prompt migration specification without rewriting the prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
sourceYes
targetYes
source_pathNo
source_roleNounknown
source_platformNo
target_platformNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does state that the tool does not rewrite the prompt, which is a meaningful behavioral constraint. However, it does not disclose what the output specification contains, whether any side effects occur, or whether the tool performs validation or analysis. The description is honest about the non-mutating nature but lacks depth about the actual behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is concise and front-loads the core purpose. It avoids redundancy and does not waste words. However, it is so brief that it sacrifices useful detail, so it earns a 4 rather than a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 7 parameters, no annotations, and an output schema, but the description does not explain what the output specification looks like, what the required parameters represent, or how this tool fits into the migration workflow. The output schema exists but the description should still clarify the tool's role relative to siblings like generate_migration_plan. The description is too thin for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the 7 parameters, but it does not explain any of them. The description mentions 'prompt migration specification' but does not clarify what 'source', 'target', or 'prompt' mean in this context, nor the optional parameters like source_path, source_role, or platform fields. The agent must infer parameter semantics entirely from names and the schema, which is insufficient for a tool with 7 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('prepare') and resource ('prompt migration specification') and adds a key constraint ('without rewriting the prompt'). This distinguishes it from sibling tools like generate_migration_plan or run_migration_eval, though it doesn't explicitly name an alternative. The phrase 'without rewriting the prompt' clarifies the tool's non-destructive intent, which is helpful for an agent deciding between this and a migration execution tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for preparing a specification rather than executing a migration, which gives some context. However, it does not explicitly state when to use this tool versus siblings like generate_migration_plan or prepare_invocation_migration, nor does it mention any prerequisites or follow-up steps. The 'without rewriting' phrase hints at a non-execution role, but the guidance is not explicit enough for an agent to confidently choose between this and similar planning tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_registry_updatePropose Registry UpdateB

Plan a registry update from supplied evidence without editing canonical files.

ParametersJSON Schema
NameRequiredDescriptionDefault
researchYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a key behavioral trait: it does not edit canonical files. Yet it does not explain what the tool does with the plan, whether it persists anything, or if it has other side effects, so transparency is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence that front-loads the verb 'Plan' and the critical constraint 'without editing canonical files'. There is no redundant wording or filler; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition is too thin for a tool that accepts a complex nested parameter and sits among many similarly named migration-planning siblings. It lacks parameter details, usage context, and relationships to alternatives, leaving an agent without enough information to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the only parameter 'research' is an open object with no documentation. The description's phrase 'supplied evidence' hints at what the parameter should contain, but it does not specify structure, required subfields, or expected format, leaving agents to guess how to populate it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Plan' with a clear resource 'registry update' and adds a crucial boundary 'without editing canonical files' that distinguishes it from mutation tools. While it does not explicitly name sibling alternatives, the non-editing qualifier sharply narrows its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'from supplied evidence' implies the tool should be used when evidence is available and a planning step is needed, providing some contextual guidance. However, it does not mention any alternatives or conditions for when not to use this tool, leaving the choice among the many sibling tools ambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_live_pricingQuery Live PricingB

Opt in to an OpenRouter pricing inquiry; canonical registry facts are unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
identifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that the tool is read-only regarding registry facts, which is helpful, but it does not disclose that it makes an external network call to OpenRouter, potential latency, failure modes, or the need for network access. This is a significant omission for a live inquiry tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no waste, but it is concise to the point of under-specification. It front-loads the key purpose and the non-mutating caveat, which is efficient. However, brevity sacrifices critical behavioral and parameter information, so while the structure is concise, the content is incomplete.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex (network call, live data) and has an output schema, but the description is minimal. It adds only the 'fact unchanged' caveat, but it omits the source of the data, the meaning of 'opt in', the expected output structure, and any prerequisites. Given the lack of annotations and 0% schema coverage, the description is not adequate for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it provides no parameter semantics. The schema defines 'identifier' and 'timeout' with no description, and the tool description does not explain what the identifier refers to (e.g., model ID or pricing ID) or what the timeout controls. This is a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific action ('opt in to an OpenRouter pricing inquiry') and clarifies that it does not modify canonical registry facts, which distinguishes its purpose from tools that update the registry. However, it does not explicitly name sibling alternatives that might be confused with it, and the term 'opt in' is somewhat idiosyncratic.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use this tool: when live pricing is needed, and it explicitly says that canonical registry facts are unchanged, which signals when not to use it (when updates are intended). However, it does not explicitly mention alternative tools or provide explicit exclusions, so there is some gap in guiding the agent away from similar inquiry tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recommend_modelsRecommend ModelsC

Hard-filter and deterministically rank compatible registry models.

ParametersJSON Schema
NameRequiredDescriptionDefault
regionNo
platformNo
providerNo
source_modelNo
migration_goalNobalanced
source_platformNo
application_pathNo
include_live_pricingNo
required_capabilitiesNo
minimum_context_windowNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does reveal meaningful traits: 'hard-filter' implies exclusion of incompatible models and 'deterministically rank' implies reproducible ordering. However, it does not clarify what 'compatible' means, how ranking works, or whether the operation is read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no filler, and the key behavioral terms are front-loaded. It is appropriately concise, though it sacrifices useful guidance for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 optional parameters, no annotations, and no parameter descriptions, this one-sentence description is insufficient for an agent to call the tool correctly. It does not explain filtering criteria, ranking rationale, or how parameters interact, and the presence of an output schema does not compensate for missing input semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no parameter-level meaning. Some parameter names like region, platform, and provider are self-explanatory, but others such as migration_goal, application_path, and source_platform remain ambiguous with no description support.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action—'hard-filter and deterministically rank'—and a clear resource, 'compatible registry models.' It conveys the tool's core function, though it does not explicitly distinguish it from siblings like compare_models or resolve_model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as compare_models, resolve_model, or get_model_profile. There are no usage scenarios, prerequisites, or exclusions, leaving the agent to infer appropriateness from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_blocker_decisionRecord Blocker DecisionA

Record the user's decision for one blocker; decisions are durable.

Ids must come from get_blocker_resolutions. Accept options REQUIRE the user's own free-text rationale and are never a default. Retarget/correction decisions update the run identity immediately; redesign decisions inject a required task on every regeneration; accept decisions downgrade the blocker to a reported accepted risk. The result lists the remaining blockers and the exact next step.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes
option_idYes
rationaleNo
blocker_idYes
decided_onNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses side effects clearly: 'decisions are durable', 'Retarget/correction decisions update the run identity immediately', 'redesign decisions inject a required task on every regeneration', and 'accept decisions downgrade the blocker to a reported accepted risk.' This is detailed and honest about consequences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it starts with the core purpose, then the prerequisite, then the behavioral details, and ends with the result. Each sentence earns its place; there is no redundant fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return details are covered. The description explains the most important aspects (source of IDs, rationale requirement, side effects, and result), but with 6 parameters and 0% schema coverage, it fails to explain run_dir, decided_on, and now. These are significant for correct invocation, so the description is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that blocker_id and option_id come from get_blocker_resolutions, and that rationale is required for accept options. However, it does not explain run_dir, decided_on, or now, leaving a gap for agents on how to populate these parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Record') and resource ('user's decision for one blocker'), and specifies scope ('one blocker'). It distinguishes from sibling tools like record_change_decision by focusing on blocker-specific decisions and mentioning the prerequisite 'Ids must come from get_blocker_resolutions'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidelines: 'Ids must come from get_blocker_resolutions' and 'Accept options REQUIRE the user's own free-text rationale and are never a default.' It also explains the different decision types and their effects, which helps the agent understand when to use this tool. However, it does not explicitly mention alternatives or when not to use it, though the name itself implies the scope.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_change_decisionRecord Change DecisionA

Record the user's accept/reject for one annotated change; durable.

change_id comes from get_change_review. Every decision deterministically regenerates the deliverable from the original, the as-submitted content, and all live rejections; if that cannot be done deterministically the decision is refused (NOT recorded) with the reason. The result lists the change ids still pending for the file.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
run_dirYes
decisionYes
change_idYes
decided_onNo
source_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses the side effect (deterministically regenerating the deliverable), the failure condition (refusal if regeneration is not deterministic), and the output (lists pending change ids). It also explicitly states the action is 'durable' (persistent). These are significant behavioral traits beyond what a schema would show.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the core purpose, and every sentence provides distinct value. It avoids repetition and is efficiently structured, with technical details placed after the primary action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the description covers the core decision logic and mentions the output, it omits explanations for required parameters like 'run_dir' and 'source_path'. Since there is an output schema, return values are likely covered, but the missing parameter semantics make the description incomplete for a tool with 6 parameters and no schema descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the source of 'change_id' and implies the values for 'decision' (accept/reject), but it does not explain 'run_dir', 'source_path', 'note', or 'decided_on'. These required parameters are left undefined, leaving significant gaps for an agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Record') and the resource ('the user's accept/reject for one annotated change'), and notes it is 'durable'. The singular 'one annotated change' differentiates it from the plural sibling tool 'record_change_decisions', so an agent can distinguish it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage context by stating that 'change_id comes from get_change_review', which tells the agent where to obtain a required parameter. It also implicitly differentiates from the plural sibling by emphasizing 'one annotated change', but it does not explicitly name alternatives or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_change_decisionsRecord Change DecisionsA

Record several review decisions in one call, with per-decision results.

Each item: {"source_path", "change_id", "decision": "accepted"|"rejected", "note"?}. Decisions apply in order (each regeneration reflects every earlier rejection); a refused decision is reported and the rest continue.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_dirYes
decisionsYes
decided_onNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses meaningful traits: decisions are applied in order, each regeneration reflects earlier rejections, refused decisions are reported, and the remaining decisions continue. This goes beyond the input schema and helps an agent predict side effects, though it could say more about persistence or validation failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: purpose first, then item shape, then behavioral rules. Every sentence earns its place, and there is no redundant restatement of the tool name or input schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch mutation tool with no annotations and 0% schema coverage, this description covers the critical semantics: item structure, ordering, refusal behavior, and per-decision results. The main gaps are the unexplained run_dir and decided_on parameters and the exact result shape, though the presence of an output schema reduces the severity of the latter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds real meaning for the decisions array item: source_path, change_id, decision with accepted/rejected values, and an optional note. However, schema description coverage is 0%, and the required run_dir parameter plus decided_on are not explained at all. The description compensates for the item schema but leaves two top-level parameters underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Record several review decisions in one call.' It clearly signals batch semantics and per-decision results, which distinguishes it from the singular sibling record_change_decision without needing to open schemas. The item format with decision enum values adds further precision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'several review decisions in one call' gives clear context for when to use this tool: when batching multiple decisions. It also communicates that decisions apply in order and that refused decisions are isolated. It does not explicitly name the singular alternative or state when not to use it, but the batch framing is strong enough guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_observationRecord ObservationA

Record the target's observed behavior for one plan unknown (run-scoped).

subject is the unknown's id (e.g. from a probe script's docstring or the worklist unknowns); outcome is what the target actually did and evidence the printed result, request id, or URL — both from the USER's own run. It closes that unknown in this run's plan and renders in the report; it never changes the registry (promotion is a separate propose_registry_update).

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
outcomeYes
run_dirYes
subjectYes
evidenceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the action is run-scoped, closes the unknown in the plan, renders in the report, and does not alter the registry. These are key behavioral side effects. It could also mention error conditions or prerequisites (e.g., existence of a run plan) but covers the most critical traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with a one-sentence purpose statement followed by a param explanation paragraph. It is reasonably concise and front-loaded, with every sentence contributing. Slightly dense but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 params, no annotations, and no schema descriptions, the description covers the core action, three of five parameters, and the most important side effect (no registry change). It omits run_dir/now semantics and does not mention any prerequisites or failure modes, but the presence of an output schema lessens the need to describe return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It meaningfully explains subject ('unknown's id'), outcome ('what the target actually did'), and evidence ('printed result, request id, or URL'), which is valuable. However, run_dir and now are left unexplained, so not all parameters get adequate semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Record the target's observed behavior for one plan unknown'. It also distinguishes from siblings by explicitly noting it 'never changes the registry' and that promotion is a separate propose_registry_update, making it easy for an agent to differentiate this from other record_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use context: 'for one plan unknown (run-scoped)' and from the USER's own run. It also gives a direct when-not-to-use signal: 'it never changes the registry (promotion is a separate propose_registry_update)', which routes the agent to the correct alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_validation_dispositionRecord Validation DispositionA

Record how this run's migration was validated; durable per run.

Record AFTER the first finalize produced the validation deliverables: byok_evaluation needs evaluation_run_path (the eval run for the current output/migration-manifest.yaml); generated_tests needs outcome "run_passed" plus the test summary line from the USER's run; an explicit accept REQUIRES the user's own rationale and is never a default. Then finalize again.

ParametersJSON Schema
NameRequiredDescriptionDefault
methodYes
outcomeNo
run_dirYes
rationaleNo
decided_onNo
outcome_summaryNo
evaluation_run_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that the record is durable, that it must happen after a specific step, that finalize must be run again afterward, and that an accept is never a default. This is meaningful behavioral context, though it stops short of describing idempotency or overwrite behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the purpose statement comes first, followed by crisp conditional requirements. Every sentence earns its place, and the line breaks make the procedural constraints easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential workflow: when to record, what each method requires, and the need to finalize again. Because an output schema exists, return-value documentation is unnecessary. It is slightly incomplete regarding how optional fields like outcome_summary and decided_on relate to each method, but it is sufficient for correct main-path invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add real semantics for evaluation_run_path, outcome, and rationale by tying them to specific method values, but it leaves run_dir, method, outcome_summary, and decided_on mostly implicit, so the compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-object pair: 'Record how this run's migration was validated' plus the qualifier 'durable per run.' This clearly distinguishes the tool from sibling record tools like record_change_decisions or record_observation by scoping it specifically to migration-validation disposition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear temporal context: 'Record AFTER the first finalize produced the validation deliverables' and then says to 'finalize again,' which tells the agent where in the workflow this tool belongs. It also gives per-method conditions, but it does not mention alternative tools or explicit when-not-to-use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_modelResolve ModelA

Match an identifier against the local registry, tolerating vague input.

Check here before researching a model: prefixes, version suffixes, and vague platform names normalize deterministically. status is resolved, needs_confirmation (show candidates to the user), or not_found.

ParametersJSON Schema
NameRequiredDescriptionDefault
endpointNo
platformNo
identifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden, and it does well by naming deterministic normalization, the three status outcomes, and the instruction to show `candidates` on `needs_confirmation`. It does not discuss side effects or permissions, but this tool appears to be a safe read-style lookup.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The core purpose is front-loaded, the usage guidance is packed into the second sentence, and the status contract is summarized compactly at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage timing, and statuses well, but the unexplained `endpoint` and `platform` parameters create a real gap for agents deciding whether to provide them. The output schema can handle return values, but it cannot compensate for missing parameter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the three parameters. It only alludes to vague platform names and version suffixes, but never explains `endpoint` or `platform`, leaving two of three parameters semantically undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource pairing: 'Match an identifier against the local registry, tolerating vague input.' It clearly distinguishes this from researching, comparing, or lifecycle-checking models, and the 'Check here before researching a model' line reinforces its role as a lookup/normalization step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Check here before researching a model' is an explicit usage directive, telling the agent this tool is the first stop for identifier resolution. It does not name specific sibling alternatives or state when not to use it, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_migration_evalRun Migration EvalC

Run source and target with user-owned local provider credentials.

ParametersJSON Schema
NameRequiredDescriptionDefault
suiteYes
source_configYes
target_configYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does reveal that the tool uses 'user-owned local provider credentials,' which is useful, but it does not say whether this operation executes provider calls, has side effects, incurs costs, requires special permissions, or how it behaves on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one tight sentence with no waste and front-loads the main action. However, its brevity crosses from concise into under-specified, so it does not quite earn a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool accepts three complex nested objects and has no annotations, so the description leaves too much to inference. The presence of an output schema covers return values, but the lack of guidance on configuration structure, execution semantics, or credentials limits completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by clarifying parameters. It does not explain what sits inside suite, source_config, or target_config, nor how they relate to one another. The parameter names alone are insufficient for three complex nested objects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Run' and names the object 'source and target,' but it does not clarify what 'source and target' actually are (models? configs? prompts?). The title 'Run Migration Eval' adds context, but the description alone does not clearly differentiate this tool from siblings like generate_eval_suite or compare_outputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, prerequisites, or alternatives. The only implicit hint is that this runs an evaluation with user-owned local provider credentials, but it does not explain when an agent should choose this over the many related sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scaffold_evaluationScaffold EvaluationB

Draft a BYOK evaluation suite from the application's own sample inputs.

Cases come from files under sample/fixture/example/evaluation folders, are marked DRAFT, and the suite (output/evaluation/) is bound to the current plan. Show the drafts to the user before anything runs; the toolkit never executes the evaluation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that cases come from specific folders, are marked DRAFT, the suite is bound to the current plan, drafts are shown before running, and the toolkit never executes the evaluation. This covers key side effects and non-actions, though it does not detail file overwrite behavior or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose. It packs in essential behavioral details efficiently without redundancy. A slightly more structured separation of purpose vs. behavior could improve readability, but it is concise and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, return format is covered, but the tool has 2 parameters with zero schema descriptions. The description fails to explain what 'now' and 'run_dir' are, nor does it mention prerequisites or the exact relationship between the sample folders and the parameters. This is a significant gap for a tool that requires a run_dir.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description does not explain the 'now' or 'run_dir' parameters at all. It mentions sample inputs and folders but does not map them to the schema, leaving the agent without any guidance on parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Draft' and the resource 'BYOK evaluation suite from the application's own sample inputs'. It distinguishes from sibling tools like 'generate_eval_suite' by emphasizing the drafting (non-executing) nature and the source of cases, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for initial drafting ('Draft', 'before anything runs') but does not explicitly contrast it with alternatives like 'generate_eval_suite' or 'run_migration_eval'. It provides context (show drafts before execution) but no explicit when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_applicationScan ApplicationA

Scan a local Python application into a normalized coupling inventory.

prompt_sources names prompt files automatic discovery missed (relative to the application root).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
prompt_sourcesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool scans an application into an inventory but does not mention whether the operation is read-only, whether it has side effects, or any prerequisites or constraints. It also does not describe the return format beyond the output schema. The description only states the purpose, not the behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no fluff. The main purpose is front-loaded, and the optional parameter explanation is concise and relevant. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core action and explains the one non-obvious parameter. Since an output schema exists, the return values are already defined. However, it does not mention any prerequisites or behavioral nuances, but for a scanning tool with an output schema, it is sufficiently complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate. It does explain the purpose of prompt_sources (names prompt files automatic discovery missed, relative to the application root), adding meaning beyond the schema. However, the 'path' parameter is not explicitly explained; its meaning is only implied by context. Thus, the description partially compensates for the lack of schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans a local Python application into a normalized coupling inventory, with a specific verb (scan), resource (local Python application), and outcome. This clearly distinguishes it from sibling tools focused on research, migration, or validation. No ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly mention when to use this tool versus alternatives, nor does it list any exclusions. The purpose is clear, so an agent could infer its usage, but there is no explicit guidance on when to choose it over other scanning or analysis tools. This is implied usage rather than explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_migrationStart MigrationA

Start a guided migration run; the preferred entry point for a full migration.

Both models match registry-first; if either needs confirmation, nothing is written and candidates must be shown to the user. Likewise, when prompt consumers exist but no prompt source resolved, prompt_candidates lists candidate files to confirm with the user: retry with prompt_sources, or defer_prompt_candidates=true to decide on the live run with add_prompt_sources. Otherwise the run workspace is created, a bounded research request is written only when knowledge is missing or stale, and next_steps says exactly what to call next. strict=true (production runs) turns unknown evidence URLs, missing invocation facts, incomplete prompt coverage, consistency findings, and a missing validation disposition into blockers/violations instead of warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofNo
run_idNo
sourceYes
strictNo
targetYes
researchNoauto
output_dirNo
prompt_sourcesNo
source_endpointNo
source_platformNo
target_endpointNo
target_platformNo
application_pathYes
defer_prompt_candidatesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses key behavioral traits: conditional early exits (nothing written and candidates shown), handling of prompt candidates, creation of a run workspace, bounded research request only when needed, and the effect of strict=true on error handling. It also implies write behavior by creating a workspace and writing research requests, which is useful. The description is detailed about side effects and conditional outcomes, going beyond a simple 'starts a migration' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately long but information-dense. It front-loads the core purpose and then explains conditional behaviors. Each sentence adds value. It could be slightly more concise, but the complexity of the tool justifies the length. No fluff or redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 14 parameters, no annotations, and an output schema present, the description covers the essential flow, conditional paths, and strict mode implications. It mentions 'next_steps' which aligns with the output schema's likely field. Given the tool's complexity, the description provides enough context for an agent to invoke it correctly, including what to do in edge cases. The presence of an output schema reduces the need to describe return values, and the description covers the main decision points.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain several parameters: strict=true behavior, defer_prompt_candidates=true, prompt_sources (indirectly via retry with prompt_sources), and research (bounded research request). It does not explain all 14 parameters, but it covers the most important ones. Given the low coverage, this is a strong effort, but it could be more comprehensive for parameters like source_platform and target_endpoint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it starts a guided migration run and is the preferred entry point for a full migration, distinguishing it from many sibling tools that cover specific sub-steps. It does not explicitly name a single sibling alternative, but the phrase 'preferred entry point' sets it apart. It could be more explicit about what it is not for, but the purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides some guidance on when to use this tool (as the preferred entry point) and covers conditional flows (e.g., when models need confirmation, when prompt consumers exist). However, it does not explicitly state when NOT to use it or name alternative tools for specific scenarios. It mentions retrying with prompt_sources but not which sibling to use instead. The guidance is embedded but not direct.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_adaptationsSubmit AdaptationsA

Submit several adaptations in one call, with per-item accept/reject.

Each item: {"kind": "prompt"|"file", "source_path", "content", "rationale", "changes"?, "new_file"?, "unchanged"?, "allow_restructure"?, "guidance_dispositions"?, "annotated_changes"?, "default_disposition"?, "default_disposition_note"?} — the same rules as the single-submission tools. Items validate independently (never all-or-nothing) and the accepted subset is applied in one locked, atomic write to changes.yaml.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNo
run_dirYes
submissionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses per-item accept/reject semantics, independent validation ('never all-or-nothing'), and an atomic locked write to changes.yaml. It doesn't cover error handling or post-conditions, but the core behavior is clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, item schema, and behavioral semantics. The field list is concise with '?' notation for optional fields, and there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for a batch wrapper: it defines the item shape, how acceptance works, and the atomic write behavior. The main gap is that the exact accept/reject signal mechanism (e.g., how default_disposition maps to acceptance) is deferred to the single-submission tools, and run_dir/now are not explained. Still, the core calling contract is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by enumerating the per-item fields with optionality markers and pointing to single-submission tools for their semantics. It does not describe run_dir or now, but the essential item structure is fully specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Submit several adaptations in one call') with a clear resource and scope. It explicitly contrasts with the 'single-submission tools' among the siblings, so an agent can distinguish batch submission from submit_adapted_prompt or submit_adapted_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrasing 'several adaptations in one call' and the reference to 'the same rules as the single-submission tools' clearly imply when to use this batch tool versus the single-item alternatives. It doesn't spell out exclusions or exact selection criteria, but the context is strong enough for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_adapted_fileSubmit Adapted FileA

Store one complete adapted application file (never a diff) for review.

Same rules as submit_adapted_prompt: dispose every required change in guidance_dispositions and document every edit in annotated_changes (anchors match the raw file content). Deterministic checks gate the write to <run>/output/files/; the application tree is never modified. Pass unchanged=true with empty adapted_content for a reviewed no-change file. Prompt sources are rejected here; use submit_adapted_prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
changesNo
run_dirYes
new_fileNo
rationaleYes
unchangedNo
source_pathYes
adapted_contentYes
annotated_changesNo
default_dispositionNo
guidance_dispositionsNo
default_disposition_noteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that deterministic checks gate the write to `<run>/output/files/`, that the application tree is never modified, and that prompt sources are rejected. This provides an accurate safety and side-effect model beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and each sentence delivers a distinct, necessary rule: file type, shared rules, write location, no-modification guarantee, no-change mode, and rejection of prompt sources. No filler or redundancy exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 11-parameter tool with no annotations, the description covers the main flow, critical constraints, key parameters, and the primary sibling alternative. Since an output schema exists, return-value details are unnecessary. The description is sufficient for correct invocation, with only minor reliance on the sibling tool for shared rules.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning for adapted_content (complete file, never a diff), unchanged (true with empty adapted_content), guidance_dispositions (dispose every required change), annotated_changes (document every edit with anchors matching raw content), and run_dir (write target). However, several parameters such as rationale, source_path, changes, new_file, and default_disposition remain unexplained, leaving some inference burden on the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Store') and resource ('complete adapted application file') and distinguishes itself from the sibling submit_adapted_prompt by explicitly stating that prompt sources are rejected. An agent can immediately tell this tool handles file content, not prompt sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use and when-not-to-use guidance: 'Prompt sources are rejected here; use submit_adapted_prompt'. It also explains the no-change case ('Pass unchanged=true with empty adapted_content') and forbids diffs ('never a diff'). These are concrete selection and invocation rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_adapted_promptSubmit Adapted PromptA

Validate and store one adapted prompt for the target model.

Adapt minimally: change only what a listed model difference or guidance item requires. Dispose EVERY guidance item of the task (its guidance plus shared_prompt_guidance) exactly once in guidance_dispositions: {guidance_id, disposition: applied|not_applicable|declined, note (required when declined)}. Document every edit in annotated_changes against the DECODED runtime values: {operation: edit|insert|delete| restructure, original_anchor/adapted_anchor (exact spans), why, evidence: [{kind: model_guidance|model_difference|analysis_finding| research|mechanical, url|reference}]}; every diff hunk covered, no phantom annotations. Serialization-, whitespace-, or case-only edits are rejected; structural drops need allow_restructure=true plus a recorded justification. No change needed? Pass unchanged=true with an EMPTY adapted_prompt (refused while the prompt still references the source model). Shared guidance already disposed earlier in the run is no longer owed, and default_disposition covers unlisted items as visibly defaulted records. Deliverables land under /output/prompts/.

ParametersJSON Schema
NameRequiredDescriptionDefault
changesNo
run_dirYes
rationaleYes
unchangedNo
source_pathYes
adapted_promptYes
allow_restructureNo
annotated_changesNo
default_dispositionNo
guidance_dispositionsNo
default_disposition_noteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses several behavioral traits: it rejects serialization/whitespace/case-only edits, refuses when unchanged=true but prompt still references source model, treats shared guidance already disposed as no longer owed, and writes deliverables to <run>/output/prompts/. It also requires allow_restructure for structural drops. This is substantive behavioral disclosure, though it does not cover permissions, idempotency, or failure modes beyond rejections.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph with many semicolon-separated rules. It front-loads the purpose but then packs in a large number of constraints with no visual structure (bullets, sections). Each sentence has content, but the density makes it harder to scan. It is longer than ideal, though appropriate for a complex validation workflow. Loses points for organization, not content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with 0% schema coverage and no annotations, this description covers a remarkable number of edge cases: minimal-adaptation policy, guidance disposition completeness, annotation exactness, rejection criteria, restructure permission, default disposition behavior, and output location. Since an output schema exists, the return value need not be described. Minor gaps: 'DECODED runtime values' and 'listed model difference' are assumed knowledge. Overall, it is nearly complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add meaning. It does—substantially. It specifies the exact structure and semantics for guidance_dispositions, annotated_changes, the meaning of unchanged (with empty adapted_prompt), allow_restructure, and default_disposition. However, it leaves several parameters (rationale, changes, default_disposition_note, source_path) under-explained, and some references ('DECODED runtime values') are ambiguous. Still, the description adds significant value beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Validate and store one adapted prompt for the target model,' naming a specific verb (validate/store), resource (adapted prompt), and scope (one, target model). This distinguishes it from broader siblings like submit_adaptations or submit_adapted_file, though it does not name them explicitly. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description is entirely procedural—how to adapt, dispose guidance, and annotate changes—but never states when to choose this tool over siblings like submit_adaptations, submit_adapted_file, or record_change_decisions. There is no explicit use-case guidance, prerequisites, or exclusion conditions. The implied usage is that this is the only submission point for adapted prompts, but that is not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_evidence_reviewValidate Evidence ReviewC

Check that an independent evidence review actually covers the research artifact.

ParametersJSON Schema
NameRequiredDescriptionDefault
reviewYes
researchYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. 'Check' implies a read-only comparison, but the description does not state whether the tool mutates state, how coverage is determined, what happens on failure, or what the output represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with an active verb and no filler. It is appropriately concise for a straightforward validation operation, though the brevity leaves behavioral and contextual details to other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two opaque nested objects and no annotations, the description is too thin. It does not explain what 'covers' means, what a valid evidence review looks like, or how the validation result should be interpreted, leaving the agent without enough context to invoke the tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only generic object types with 0% description coverage, so the description must compensate. It adds a useful semantic mapping: review corresponds to the evidence review and research to the research artifact. However, it does not explain required properties, nesting, or the structural relationship between the two objects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('check') applied to a specific resource ('independent evidence review') against another resource ('research artifact'). It conveys the tool's role in verifying coverage, though it does not explicitly distinguish it from siblings like validate_research_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as validate_research_result or build_research_consensus. No workflow context, prerequisites, or exclusions are provided; the description only implies a verification scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_promptValidate PromptC

Statically validate prompt assumptions against target registry facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
targetYes
source_pathNo
target_endpointNo
target_platformNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It says the validation is 'static' but does not explain side effects, failure modes, what happens when assumptions conflict with registry facts, or whether this is a read-only operation. The description provides only a high-level intent, not enough behavioral transparency for a tool with no annotation safety hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is compact and readable. It loses a point because 'statically validate' is slightly jargon-heavy and the sentence carries insufficient detail, but as a standalone statement it is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has five parameters, zero annotation coverage, and no parameter documentation in the schema. Although an output schema exists, the description still fails to convey what inputs are expected beyond the two required ones, when to use the tool, or what the validation guarantees. This is inadequate for correct invocation in an agentic context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the five parameters. It only hints at 'prompt assumptions' and 'target registry facts,' which loosely maps to prompt and target, but it never explains source_path, target_endpoint, or target_platform. The optional parameters remain completely undocumented, leaving an agent without enough meaning to populate them correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('validate'), a specific object ('prompt assumptions'), and a reference point ('target registry facts'). The qualifier 'statically' adds mode-of-operation detail. It does not explicitly name sibling tools, but the action is specific enough to separate it from broader analysis tools like analyze_prompt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to choose this tool over siblings such as analyze_prompt, validate_evidence_review, or run_migration_eval. There are no exclusions, prerequisites, or alternative conditions. Usage context is only weakly implied by the verb 'validate'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_research_artifactValidate Research ArtifactA

Validate one scope's research artifacts directly from the run workspace.

Reads request.yaml, research/<scope>.yaml, and (when present) review/<scope>.yaml and runs the deterministic scope/policy and review-integrity gates; nothing is resent through the payload. scope is source, target, or pair.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeYes
run_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does a solid job: it names the exact input files, notes the optional review file, states the gates are deterministic, and clarifies that nothing is resent through the payload. It could add explicit no-side-effect or permission details, but the read-and-validate behavior is clearly conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: the first sentence states the purpose, the second details the exact behavior and files, and the third pins down the scope values. Every sentence earns its place, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists, return-value details are unnecessary. The description covers input files, gates, scope values, and the no-resend behavior, which is enough for an agent to invoke it correctly. It could be more explicit about run_dir semantics, but overall the essentials are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning for scope by listing valid values and showing how it interpolates into file paths, and it ties run_dir to the run workspace. However, it never explicitly names or describes run_dir as the workspace directory parameter, leaving some inference required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Validate') and resource ('one scope's research artifacts'), then enumerates the exact files and gates involved. It also names the valid scope values, which separates it from sibling validation tools such as validate_research_result and validate_evidence_review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context ('directly from the run workspace') but no explicit guidance on when to choose this tool over alternatives like validate_research_result or validate_evidence_review. There are no exclusions, prerequisites, or alternative tool mentions, so an agent must infer the usage boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_research_resultValidate Research ResultC

Run the deterministic scope/policy gate over one research artifact.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes
researchYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavior disclosure, yet it only reveals that the operation is deterministic and gate-like. It does not state whether the tool is read-only, what happens when a result fails the gate, or what side effects or return behavior to expect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It delivers the core action and resource in the fewest possible words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

At 0% parameter documentation and no annotations, a one-sentence description is insufficient for a tool with two opaque nested objects and a large sibling set. The description needs to clarify the role of 'request,' the nature of the gate result, and when this tool applies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the prose only weakly hints at the 'research' parameter via 'one research artifact.' The 'request' parameter is entirely unexplained, and the nested object structures have no semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete operation ('run a deterministic scope/policy gate') and a specific resource ('one research artifact'), making the core function understandable. It is somewhat generic relative to sibling validate tools, but the artifact-and-gate framing distinguishes it enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus validate_prompt, validate_evidence_review, or other validation siblings. The description states what the tool does but not under what conditions an agent should choose it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 15 tool updatesv1.6.0
    • Addedadd_prompt_sources
    • Addedconfirm_prompt_consumer
    • Addedconfirm_unaffected
    • Addedget_change_review
    • Addedget_run_status
    • Addedrecord_change_decision
    • Addedrecord_change_decisions
    • Addedrecord_observation
    • Addedrecord_validation_disposition
    • Addedscaffold_evaluation
    • Changedstart_migration2 fields changed
      • addedInput schema / properties / defer_prompt_candidates
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
      • addedInput schema / properties / strict
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
    • Addedsubmit_adaptations
    • Changedsubmit_adapted_file9 fields changed
      • addedInput schema / properties / annotated_changes
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "additionalProperties": true,
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedInput schema / properties / changes / anyOf
        Added value: +[
        +  {
        +    "items": {
        +      "type": "string"
        +    },
        +    "type": "array"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / changes / default
        Added value: +null
      • removedInput schema / properties / changes / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / changes / type
        Removed value: -"array"
      • addedInput schema / properties / default_disposition
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "applied",
        +        "not_applicable",
        +        "declined"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedInput schema / properties / default_disposition_note
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
      • addedInput schema / properties / guidance_dispositions
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "additionalProperties": {
        +          "type": "string"
        +        },
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • changedInput schema / required
        Previous value: -[
        -  "run_dir",
        -  "source_path",
        -  "adapted_content",
        -  "rationale",
        -  "changes"
        -]New value: +[
        +  "run_dir",
        +  "source_path",
        +  "adapted_content",
        +  "rationale"
        +]
    • Changedsubmit_adapted_prompt4 fields changed
      • addedInput schema / properties / annotated_changes
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "additionalProperties": true,
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedInput schema / properties / default_disposition
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "applied",
        +        "not_applicable",
        +        "declined"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedInput schema / properties / default_disposition_note
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
      • addedInput schema / properties / guidance_dispositions
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "additionalProperties": {
        +          "type": "string"
        +        },
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Addedvalidate_research_artifact
  2. 8 tool updatesv1.3.0
    • Changedgenerate_migration_plan1 field changed
      • addedInput schema / properties / prompt_sources
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Changedgenerate_migration_report1 field changed
      • addedInput schema / properties / prompt_sources
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Addedget_blocker_resolutions
    • Addedrecord_blocker_decision
    • Changedscan_application1 field changed
      • addedInput schema / properties / prompt_sources
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Changedstart_migration1 field changed
      • addedInput schema / properties / prompt_sources
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Changedsubmit_adapted_file1 field changed
      • addedInput schema / properties / unchanged
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
    • Changedsubmit_adapted_prompt2 fields changed
      • addedInput schema / properties / allow_restructure
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
      • addedInput schema / properties / unchanged
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
  3. 9 tool updatesv1.2.0
    • Addedfinalize_migration
    • Changedgenerate_migration_plan2 fields changed
      • addedInput schema / properties / source_endpoint
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedInput schema / properties / target_endpoint
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Changedgenerate_migration_report2 fields changed
      • addedInput schema / properties / source_endpoint
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedInput schema / properties / target_endpoint
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Addedget_research_prompts
    • Addedlist_adaptation_tasks
    • Addedstart_migration
    • Addedsubmit_adapted_file
    • Addedsubmit_adapted_prompt
    • Changedvalidate_prompt1 field changed
      • addedInput schema / properties / target_endpoint
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
  4. 27 tool updatesv1.1.0
    • First observedanalyze_invocation
    • First observedanalyze_prompt
    • First observedanalyze_regressions
    • First observedbuild_research_consensus
    • First observedbuild_session_registry
    • First observedcheck_model_lifecycle
    • First observedcompare_models
    • First observedcompare_outputs
    • First observedcreate_migration_research_request
    • First observedestimate_migration_cost
    • First observedgenerate_eval_suite
    • First observedgenerate_migration_plan
    • First observedgenerate_migration_report
    • First observedgenerate_session_migration_plan
    • First observedget_model_profile
    • First observedoptimize_migration
    • First observedprepare_invocation_migration
    • First observedprepare_prompt_migration
    • First observedpropose_registry_update
    • First observedquery_live_pricing
    • First observedrecommend_models
    • First observedresolve_model
    • First observedrun_migration_eval
    • First observedscan_application
    • First observedvalidate_evidence_review
    • First observedvalidate_prompt
    • First observedvalidate_research_result

TDQS

B3/5.0

Scored across 47 tools

Disambiguation4/5

Most tools map to distinct workflow stages and the detailed descriptions make boundaries clear. A few near-twins remain, notably validate_research_artifact vs. validate_research_result and generate_migration_plan vs. generate_session_migration_plan, where an agent could plausibly pick the wrong one.

Naming Consistency5/5

All tools use imperative lower_snake_case verb_noun names with strong prefixes like get_, record_, validate_, analyze_, generate_, and submit_. Even batch/single pairs like record_change_decision(s) and submit_adaptations follow the same pattern; the minor eval/evaluation abbreviation is not enough to break consistency.

Tool Count2/5

47 tools is far beyond the typical well-scoped 3-15 tool surface and above the 25+ threshold. The count burdens tool selection; several near-duplicate batch/single and workspace/direct validators could be consolidated without losing capability.

Completeness5/5

The domain is covered end-to-end: scan, research, plan, prompt/file adaptation, change review, blockers, finalization, evaluation, regression analysis, and registry/cost helpers all exist. There are no obvious dead ends; even rare needs such as no-change files, validation disposition, and batch decisions are handled.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables coding agents to perform safe, project-wide Python refactoring (rename, move, extract, inline, change signature, organize imports, etc.) with a dry-run safety contract and LSP-coordinate addressing.
    15
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI coding agents to inspect project dependency versions, resolve version constraints, and audit Python code for deprecated or incompatible APIs using authoritative evidence and local AST analysis.
    3
    20 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to perform read-only source investigation via a local SQLite index, offering status, search, backlog, context, impact, and drift tools for TypeScript, JavaScript, and Java codebases.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides coding agents with a Python repository's call graph and impact analysis, enabling them to see callers, callees, and affected tests before making changes.
    1
    MIT