Skip to main content
Glama

universal-os

Skills, tools, an MCP server and a shared knowledge base that let an AI coding agent reverse engineer an existing operating system from its ISO and rebuild it — ReactOS style — component by component, until apps written for the original OS run on the rebuild.

Works with Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot, OpenCode, or anything that reads AGENTS.md or speaks MCP. Point the agent at a Windows ISO (any version: XP, 7, 10, 11, Server), it fingerprints the image, boots it in a headless QEMU sandbox, surveys what the OS actually exposes, writes behavioral specs, scaffolds ReactOS-style C components, builds them, boots the rebuild, and tests real Windows apps on it. Every session leaves field notes the next agent starts from.

   ISO ──► fingerprint ──► sandbox VM ──► survey (API surface, traces)
                                            │
            behavioral specs ◄──────────────┘
                   │
            scaffold + build (ReactOS-style C, .spec exports)
                   │
            rebuilt image ──► boot in VM ──► test real apps ──► behavior_diff
                                                                    │
                                                        field note for the next agent

Install

Pick your agent. Each gets the same skills, the universal-os MCP server, and the uos CLI. Everything below installs straight from this repo — no PyPI release needed yet.

Agent

Install

Claude Code

claude mcp add universal-os -- uvx --from git+https://github.com/Jenischristan/universal-os uos-mcp — or clone the repo and open it in Claude Code (.mcp.json works out of the box)

Codex / Cursor / VS Code

Clone the repo and start the agent inside it; configs in .codex/config.toml, .cursor/mcp.json, .vscode/mcp.json are self-contained (they call uvx --from git+... uos-mcp)

Skills only (any agent)

npx skills add github.com/Jenischristan/universal-os

Anything else

git clone https://github.com/Jenischristan/universal-os and start your agent inside it

The CLI, anywhere:

uv tool install git+https://github.com/Jenischristan/universal-os   # or: pip install git+https://github.com/Jenischristan/universal-os

Both entry points (uos, uos-mcp) land on your PATH. There is no universal-os-mcp command — the MCP server executable is uos-mcp (uos mcp is the same thing).

You need Python 3.10+, and QEMU for the VM parts (git, cmake, ninja and a MinGW toolchain when you reach the build phase):

sudo apt-get install -y qemu-system-x86 qemu-utils   # Debian/Ubuntu
brew install qemu                                    # macOS
winget install SoftwareFreedomConservancy.QEMU       # Windows

uos env checks all of it and prints what's missing.

Related MCP server: EdgeBox

Try it

Rebuild the core of the Windows XP SP3 ISO at ~/isos/en_win_xp_sp3.iso far enough to run notepad.exe and cmd.exe.

This is a Windows 7 ISO. Fingerprint it, boot it in a VM, dump the kernel32/user32/ntdll API surfaces, and tell me how far the ReactOS component set already covers it.

Take my rebuilt workspace, test Notepad2 on it, and add compat quirks for everything that breaks.

The agent starts with the build-any-os skill and runs the same loop every time:

  1. search the knowledge base; 2. fingerprint the ISO; 3. check the environment;

  2. set up a sandbox lab (VM, snapshots); 5. survey the reference OS (API surfaces, traces);

  3. pick the component route; 7. write behavioral specs; 8. build one working slice;

  4. boot the rebuild and test real apps; 10. write a field note for the next agent.

A knowledge base AIs write for AIs

knowledge/ holds field notes: how specific OS rebuilds actually went — the exact ISO versions, the component order that worked, the interface facts, the verification, the gotchas (symptom → cause → fix). Every agent that finishes a milestone can open a PR with its note, so the next agent starts where the last one left off.

uos kb search "windows xp ntdll"     # before you start: prior art
uos kb new --os "windows-xp" --component ntoskrnl --title "First boot" --agent "Claude"
uos kb check knowledge/os/windows-xp/....md

Contribution rules, for humans and AIs, are in RULES.md: no Microsoft binaries, no decompiled code, no media links, honest status and verification.

What's inside

Part

What it does

MCP server (uos-mcp)

39 tools over stdio: ISO inspect/fingerprint, VM lifecycle (boot, screenshot, snapshot, exec, file transfer), PE analysis, API-surface dumps, syscall-trace plans, behavior diffs, component scaffolds, compat shims, reference-repo cloning, app-compat testing, knowledge base

uos CLI

The same tools from a shell — for agents without MCP and for humans

skills/build-any-os

The whole loop, hard rules, and playbooks per phase (fingerprint, boot, survey, route, spec, build, verify, publish)

knowledge/

Field notes from previous rebuilds + a generated INDEX

prompts/

Ready-made phase prompts for any MCP client

examples/

A worked example: workspace layout, first build, first app test

CI

GitHub Action running the tests, CLI smoke checks and knowledge-base lint on every push

Rules it follows (moderate level — full text in RULES.md)

  • Everything untrusted runs inside the sandbox VM, never on the host. The reference ISO is never mounted on the host; it is read with the built-in ISO parser or booted in QEMU.

  • Inspection is free; shipping is not. Inside the VM the agent may run, trace, dump and analyze anything. What may enter the rebuilt OS repo is: interface facts (names, ordinals, versions, structures), behavioral specs, and clean-room code. What may not: Microsoft binaries, drivers, fonts, registry hives, decompiled output, links to OS media.

  • VM networking is OFF by default; enabling it needs the user's explicit OK.

  • Snapshot before risky operations; restore to undo.

  • Keep a journal: WORKSPACE.md in the workspace. It becomes the field note at the end.

  • No activation/DRM/licensing bypass, no pirate media, ever.

License and stance

MIT. universal-os is tooling for interoperability and education — the same legal footing as ReactOS and Wine. It ships no Microsoft code, no Microsoft binaries, and does not distribute OS media. What you build with it must pass the same rules; uos kb check and CI enforce the basics, and your jurisdiction and lawyer own the rest.

Available Tools

39 tools
analyze_peA

Analyze one PE binary (headers, sections, imports, exports, version resource).

Works on a file extracted INSIDE the VM sandbox or copied out with vm_get_file. Reports interface facts only — no disassembly, nothing shippable from the binary.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose an important scope boundary ('no disassembly, nothing shippable from the binary'), which is genuinely useful behavioral context. But it says nothing about permissions, whether the target must be a running VM, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences: purpose first, then source and scope limitation. No filler, though the second paragraph packs source and limitations together without much structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers purpose, input source, and scope limits. For a single-parameter analysis tool this is nearly complete; the missing piece is any note on environment prerequisites or read-only safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'path' parameter, so the description must compensate. It does clarify the origin of the file (inside the VM or copied out via vm_get_file), which gives the agent some semantic grounding for what path means, but it omits format, relative vs absolute, and sandbox path conventions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Analyze) and resource (PE binary), then enumerates exactly what is extracted: headers, sections, imports, exports, version resource. This is precise enough to distinguish it from siblings like dump_api_surface or iso_inspect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells the agent where the input file must come from ('extracted INSIDE the VM sandbox or copied out with vm_get_file'), which is real usage context. However, it never states when to choose this over related siblings (dump_api_surface, iso_inspect) or any preconditions beyond file origin.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

behavior_diffB

Diff two API-surface dumps (reference vs rebuilt build): per-module export coverage, missing/extra exports. THE key progress metric for compatibility.

ParametersJSON Schema
NameRequiredDescriptionDefault
outNo
currentYes
baselineYes
workspaceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It discloses the computed result set (coverage, missing/extra exports) and the comparison nature implies a read-only operation, but it does not state permissions, side effects, or error behavior. This is adequate but leaves meaningful behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence with the core action front-loaded: 'Diff two API-surface dumps.' The colon-delimited output summary is efficient. The phrase 'THE key progress metric' adds useful emphasis but is slightly evaluative rather than purely operational.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter tool with 0% schema description coverage and no annotations, the description does too little on the input side. It does not explain what baseline, current, out, or workspace should contain, nor does it state any operational constraints. The output schema covers return values, but the description remains incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters. The description hints that two inputs are the reference and rebuilt dumps, which loosely maps to 'baseline' and 'current,' but it never names them or explains 'out' and 'workspace.' Two of four parameters are effectively undocumented in both the schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Diff two API-surface dumps (reference vs rebuilt build).' It also states the output content, 'per-module export coverage, missing/extra exports,' which makes the purpose intelligible without opening a schema. It does not explicitly contrast itself with sibling tools like dump_api_surface or compat_report, so it falls short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the scenario: compare a reference dump against a rebuilt build to measure compatibility progress. However, it never explicitly says when to call this instead of alternatives such as dump_api_surface or compat_report. The usage context is inferable but not clearly specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_planC

Return the concrete build plan for the workspace and check which host toolchains are present (cmake/ninja/mingw/RosBE). The AI runs the steps in src/.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolchainNorosbe
workspaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It implies a read-only operation ('Return... and check'), but never states side effects, required permissions, or whether the plan is cached/generated. The note that 'The AI runs the steps in src/' is useful workflow context but not enough disclosure for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and the workflow hint trailing. Nothing is wasted, though the phrasing 'The AI runs the steps in src/' is slightly ambiguous about who executes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but with zero schema coverage on two parameters and no annotations, the description leaves key invocation details (toolchain default/valid values, workspace semantics) unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It partially does by naming the toolchain options (cmake/ninja/mingw/RosBE), which maps to the 'toolchain' parameter, but it never explains the default 'rosbe' or the required 'workspace' parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: returning a build plan for a workspace and checking host toolchains, with concrete examples (cmake/ninja/mingw/RosBE). This is clearly distinct from VM-lifecycle and workspace-setup siblings, though it doesn't explicitly name the sibling (env_check) it overlaps with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use statement, no prerequisites, and no routing to alternatives. The overlap with env_check and workspace_status is left for the agent to infer, and there is no indication of when this should run in the build workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cleanup_workspaceA

Find (and with confirm=true delete) regenerable junk in the workspace: pycache, .pytest_cache, build/, dist/, .pyc/.pyo/*.tmp. Dry run by default - the response lists what would go. Sources, WORKSPACE.md, reports/ and third-party/ are never touched; VMs live under UOS_HOME and are unaffected. Use before archiving or sharing the workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo
workspaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the destructive action is gated behind confirm=true, defaults to a non-destructive dry run, and lists explicit exclusions that are never touched (sources, WORKSPACE.md, reports/, third-party/, VMs under UOS_HOME). This is exactly the behavioral context an agent needs before a mutating call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the operation and target globs, then default behavior, exclusions, and usage trigger. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite a 2-param schema with no annotations, the description covers the mutation gate, the safe default, the exclusion set, and the use case; an output schema exists so return-value detail is unnecessary. Nothing needed to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must supply meaning. It fully explains confirm ('with confirm=true delete', 'dry run by default') but leaves the workspace parameter's expected form (path vs. name) implicit. Partial compensation warrants the baseline-ish 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('find/delete regenerable junk in the workspace') and enumerates the exact targets (__pycache__, .pytest_cache, build/, dist/, *.pyc/*.pyo/*.tmp), which no sibling tool covers. An agent can distinguish it from VM- and analysis-oriented siblings at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use trigger ('Use before archiving or sharing the workspace') and states the dry-run-default posture, which sets the operating context. It does not name an alternative tool, but no sibling performs cleanup, so explicit routing isn't needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clone_reference_repoA

Clone a reference implementation into third-party/ (read-only reference): reactos, wine, qemu, virtio-win, or any git URL. Clone repos as needed when a component needs proven patterns or test suites.

ParametersJSON Schema
NameRequiredDescriptionDefault
destNo
repoNoreactos
workspaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully characterizes the result as a 'read-only reference' and fixes the write location (third-party/), but says nothing about network requirements, clone duration/size, or what happens if the destination already exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the action, destination, and accepted values, followed by the usage condition. Every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers action, destination, repo values, and intent. It falls slightly short given zero annotation coverage and zero schema descriptions, leaving the workspace parameter and network/overwrite behavior unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It adds real value for 'repo' by listing accepted values (none of which appear in the schema as an enum) and pins 'dest' to third-party/, but the required 'workspace' parameter is never explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Clone) plus resource (a reference implementation) and the exact destination (third-party/), and enumerates the accepted repo values (reactos, wine, qemu, virtio-win, or any git URL). No sibling tool performs a clone, so an agent can identify this unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit trigger condition: 'Clone repos as needed when a component needs proven patterns or test suites.' There is no overlapping alternative among the siblings, so no exclusions are needed, but it stops short of stating prerequisites (e.g., network availability).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compat_reportB

Aggregate all recorded app tests into a dated compatibility report (pass rate, failing labels, next-step guidance).

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the report's contents (pass rate, failing labels, next-step guidance). However it omits whether the operation is read-only, what happens when no tests have been recorded, and any workspace-scoping behavior, leaving meaningful gaps for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero waste; the action and output artifact are front-loaded and the parenthetical enumerates report contents efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required, and the description summarizes the report well. It still leaves the workspace parameter unexplained and gives no precondition about tests having been recorded, which is a gap for a tool with no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required 'workspace' parameter is undocumented in both the schema and the description. The phrase 'all recorded app tests' never clarifies whether the aggregation is scoped to the workspace argument, so the description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Aggregate') and resource ('all recorded app tests') and names the artifact produced ('dated compatibility report'), which separates it from test-running siblings like test_app_in_vm. It does not explicitly contrast itself with regression_log, the nearest sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use statement, no prerequisites (e.g. tests must already be recorded), and no named alternative such as regression_log. The aggregation timing is only weakly implied by 'recorded app tests'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dump_api_surfaceB

Dump the interface surface of a directory of reference DLLs: export names, ordinals, imports, versions -> JSON + MD. This is the rebuild's target contract.

IMPORTANT (moderate guardrails): keep the inspected copies OUTSIDE the repo (e.g. ~/inspect/); only the generated surface JSON/MD goes into workspace/api/.

ParametersJSON Schema
NameRequiredDescriptionDefault
outYes
patternNo*.dll
src_dirYes
workspaceNo
include_exeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden, and it does disclose real side-effect context: output artifacts are JSON + MD and only those land in workspace/api/, while inspected copies must stay outside the repo. It still omits whether the tool writes/deletes anything else, permission needs, or failure behavior for non-PE inputs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and output format in the first sentence, with the guardrail clearly set apart. Nearly every clause earns its place, though the arrow notation and parenthetical label are slightly terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return values need no explanation, and the guardrail adds genuine operational context. But for a five-parameter tool with zero schema coverage and no annotations, the undocumented optional parameters (pattern, include_exe, workspace) leave the definition short of what an agent needs to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so all five parameters rely on the description, which only partially compensates. It clarifies src_dir (a directory of reference DLLs) and hints at the output destination, but never explains pattern, include_exe, or workspace, leaving half the surface undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Dump) and resource (interface surface of a directory of reference DLLs) and enumerates exactly what is extracted: export names, ordinals, imports, versions, emitted as JSON + MD. The one gap is that it never distinguishes itself from the closely related analyze_pe sibling, so an agent could reasonably confuse the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Framing it as 'the rebuild's target contract' implies it belongs early in the rebuild workflow (before stub/compat-layer generation), which is useful contextual guidance. However, it never states when to prefer this over analyze_pe or iso_inspect, nor any prerequisites for calling it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

env_checkA

Check the host environment: qemu, qemu-img, git, cmake, ninja, mingw, /dev/kvm. Run once at session start; the result decides what the AI can do right now.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden, and it discloses meaningful behavioral context: this is a session-start diagnostic whose result 'decides what the AI can do right now.' The verb 'Check' plus the enumerated probes imply a read-only capability test, though it never explicitly states there are no side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The what-is-checked list is front-loaded and the operational timing note follows immediately. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. For a zero-parameter diagnostic the description covers what is checked, when to run it, and why it matters. The only minor omission is an explicit read-only/no-side-effect statement.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to document. The baseline for a zero-parameter tool is 4, and nothing in the description contradicts or undercuts that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Check) and resource (host environment), then enumerates exactly what is probed: qemu, qemu-img, git, cmake, ninja, mingw, /dev/kvm. This concrete list makes it unmistakably distinct from VM-oriented siblings like vm_status or workspace_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit timing guidance: 'Run once at session start.' That is clear when-to-use context. It does not name a when-not condition or an alternative tool, but for a session-bootstrap diagnostic there is no obvious sibling to route against.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fingerprint_windowsA

Identify the Windows version/build inside an install ISO (any version): layout markers, version resources of setup binaries, cversion.ini — with confidence.

Run this first. It decides the API-surface strategy (NT 5.x vs 6.x differ a lot). Version facts are interface facts — fine to journal; media contents stay out of the repo.

ParametersJSON Schema
NameRequiredDescriptionDefault
iso_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses what the tool inspects and that it reports 'with confidence', plus a journaling policy ('version facts ... fine to journal; media contents stay out of the repo'), but says nothing about failure behavior or return structure. Adequate, not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose first, then the 'run this first' directive, then the journaling note. The final sentence is somewhat tangential but short enough to earn its place as policy context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described. For a one-parameter read-only fingerprint tool, the description covers purpose, ordering in the workflow, detection signals and confidence. The only gap is parameter detail, which is minor for a self-evident iso_path.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single iso_path parameter. The description contextualizes it only as 'an install ISO', adding no path format, host-vs-guest location, or validation details, so it does not compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Identify the Windows version/build inside an install ISO', and enumerates the signals examined (layout markers, version resources, cversion.ini). It is clear what the tool does, though it does not explicitly contrast itself with the ISO siblings (iso_inspect, iso_verify).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Run this first. It decides the API-surface strategy (NT 5.x vs 6.x differ a lot)' gives clear when-to-use sequencing and rationale. However, it never names the alternative tools or states when NOT to use it, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gen_api_stubB

Generate .spec + stub .c for one DLL straight from an API-surface dump (stubs return E_NOTIMPL; implement from behavioral specs, highest-import first).

ParametersJSON Schema
NameRequiredDescriptionDefault
moduleYes
max_stubsNo
workspaceYes
surface_jsonYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but does disclose a key behavioral trait: stubs return E_NOTIMPL, i.e. they are non-functional placeholders. It omits file-writing side effects, whether existing files are overwritten, and any workspace requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the core action and output artifacts front-loaded; the parenthetical adds sequencing context efficiently without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, and the core generate-from-dump behavior is covered. But for a file-writing tool with no annotations and 0% parameter documentation, the gaps around workspace and max_stubs leave the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all four parameters. It only loosely maps to two of them ('one DLL' ~ module, 'API-surface dump' ~ surface_json) and says nothing about workspace or max_stubs, leaving half the surface undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Generate) and concrete resources (.spec + stub .c) scoped to 'one DLL' and sourced from 'an API-surface dump'. It clearly differs from dump_api_surface (which produces the dump) and gen_compat_layer, though it doesn't name those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical implies sequencing ('implement from behavioral specs, highest-import first') and hints this comes after a surface dump, but it never states when to choose this tool over scaffold_component or gen_compat_layer, nor any prerequisites like workspace setup.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gen_compat_layerC

Scaffold a small compat shim for one module: records per-quirk workarounds that unblock specific apps (quirks: semicolon-separated; for_apps: comma-separated).

ParametersJSON Schema
NameRequiredDescriptionDefault
moduleYes
quirksNo
for_appsNo
workspaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It implies a write/scaffold operation but says nothing about what files or directories get created, whether existing files are overwritten, required permissions, or reversibility. The output schema covers return values, but side effects remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and its purpose, with parameter encodings packed into a trailing parenthetical. Dense but no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema relieves the description of explaining return values, and the two required params are self-describing by name. However, for a mutation tool with zero annotation coverage and 0% schema description coverage, the absence of any side-effect or prerequisite detail leaves a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It helpfully documents the encoding of two params (quirks semicolon-separated, for_apps comma-separated), but module and workspace are left to inference from their names. Partial compensation for a 4-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Scaffold') and resource ('compat shim for one module') plus the payload it records (per-quirk workarounds for specific apps). It is clear on its own, but does not differentiate itself from closely related siblings like scaffold_component, gen_api_stub, or compat_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no named alternative. The phrase 'unblock specific apps' hints at the motivating scenario, but an agent must infer whether this precedes or follows compat_report, scaffold_component, or test_app_in_vm.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

init_workspaceB

Create a rebuild workspace: src/{boot,drivers,subsystems,dll,shell,include,apps}, spec/, api/, tests/, reports/, third-party/, plus WORKSPACE.md journal.

Everything the AI builds lands here. target examples: windows-xp, windows-7, windows-any (version decided after fingerprint_windows).

ParametersJSON Schema
NameRequiredDescriptionDefault
isoNo
nameNo
pathYes
targetNowindows-any
versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It does disclose real behavioral facts: it creates a filesystem tree plus a WORKSPACE.md journal, and it hints target selection is deferred. It omits what happens on an existing path (overwrite? error?), and any permission or destructive implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the concrete output layout, then the target hint. Every phrase carries information; no filler or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the directory layout is thorough. But for a 5-parameter initialization tool with zero schema documentation and no annotations, the three undocumented arguments leave a real invocation gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate. It only partially covers 'target' by giving examples (windows-xp, windows-7, windows-any) and alludes to 'version' via fingerprint_windows, but says nothing about the required 'path', nor 'iso' or 'name', leaving an agent guessing about three of five parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('create') and resource ('rebuild workspace') and enumerates the concrete directory layout and WORKSPACE.md journal it produces, so an agent knows exactly what artifact appears. It does not name or differentiate itself from siblings like workspace_status or cleanup_workspace, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'version decided after fingerprint_windows' implies a call-ordering prerequisite, which is genuine usage guidance. However it never states when this tool should be used versus alternatives (e.g. relative to workspace_status or cleanup_workspace) or whether it should precede or follow vm_create.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

iso_extractA

Extract ONE file from an install ISO to the host inspect dir — no mounting (pure ISO9660/Joliet read), e.g. member='/I386/NTOSKRNL.EX_' dest='~/inspect/xp/'.

Inspection only: keep extracted copies OUTSIDE the rebuilt-OS repo and keep only interface facts (RULES.md). Pair with analyze_pe / dump_api_surface.

ParametersJSON Schema
NameRequiredDescriptionDefault
destYes
max_mbNo
memberYes
iso_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the read mechanism (pure ISO9660/Joliet, no mount) and that output lands on the host inspect dir. It is silent on behavior the agent needs to trust: what happens when the member is absent, whether dest files are overwritten, and what the max_mb default of 64 does when exceeded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and its constraint, followed by hygiene and pairing advice. The path example earns its space; nothing is padded, though 'RULES.md' is a niche reference.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not needed, and the description covers mechanism, destination, and workflow placement. The only real gap is undocumented max_mb semantics and unspecified overwrite/error behavior for a tool that writes to the host filesystem.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it partly does via the worked example for member and dest, with iso_path implied. max_mb — the size guard with a default of 64 — is never mentioned, leaving a required-for-safety parameter entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource — extracting ONE file from an install ISO to the host inspect dir — and adds a concrete example with real member/dest values. The 'no mounting (pure ISO9660/Joliet read)' clause implicitly distinguishes it from mount-based siblings like vm_mount_iso, though it never names that alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives workflow guidance ('Inspection only', keep copies outside the rebuilt-OS repo, keep only interface facts in RULES.md) and names downstream tools (analyze_pe, dump_api_surface), but never states when to choose this over iso_inspect, iso_verify, or vm_mount_iso, all of which read/extract ISO content. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

iso_inspectA

List an install ISO's contents WITHOUT mounting it (pure ISO9660/Joliet read).

Use this before any VM work: confirms the image is readable, shows the layout (I386/ legacy vs SOURCES/ modern), and surfaces version markers. Run fingerprint_windows for the full version verdict.

ParametersJSON Schema
NameRequiredDescriptionDefault
iso_pathYes
max_entriesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does convey the important trait that nothing is mounted and the read is pure ISO9660/Joliet, implying a non-destructive inspection, but it says nothing about permissions required, side effects, or cost. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The value proposition is front-loaded in the first sentence, followed by the when-to-use context and the routing note. Three tight sentences with no filler, though the version-marker detail could be folded more economically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. For a two-parameter read-only inspection tool the description covers purpose, timing, expected findings, and a hand-off to fingerprint_windows; only the parameter meanings are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description mentions neither iso_path nor max_entries. In particular the max_entries default of 300, which governs how much of the listing is returned, is left completely unexplained, so agents cannot anticipate truncation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (list an install ISO's contents) plus the key constraint 'WITHOUT mounting it', which immediately separates it from vm_mount_iso and iso_extract. The parenthetical 'pure ISO9660/Joliet read' adds precision that an agent can use to distinguish it from sibling ISO tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this before any VM work' gives a clear precondition, and it names fingerprint_windows as the alternative for the full version verdict. It does not explicitly distinguish itself from the nearest siblings iso_verify, iso_extract, or vm_mount_iso, so routing among those is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

iso_verifyA

Validate an install ISO before use. Checks that the media exists and contains a recognizable Windows boot/install structure with no host-side mounting.

ParametersJSON Schema
NameRequiredDescriptionDefault
iso_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a meaningful behavioral trait: the check is non-invasive because it performs 'no host-side mounting.' It does not address failure behavior, error signaling, authentication, or any side effects beyond that single constraint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the action front-loaded and the non-mounting constraint following immediately. Every clause carries information; no filler or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and a single required parameter keeps the surface small. Purpose, scope, and the key non-side-effect constraint are covered; only explicit sibling routing and path expectations are missing, which are minor given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter (iso_path) at 0% schema description coverage, so the description should compensate. It clarifies that the input is an install ISO image, which gives some context, but adds nothing about accepted path forms (absolute, UNC, URL) or file expectations. Marginal value over the self-describing parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Validate an install ISO') and enumerates the checks (media exists, recognizable Windows boot/install structure), which is concrete. The clause 'with no host-side mounting' implicitly differentiates it from vm_mount_iso, but the sibling it is closest to (iso_inspect) is never named, so an agent must still infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'before use' gives a clear triggering context for when to call it, which is better than nothing. However, no alternatives are named (iso_inspect, iso_extract, vm_mount_iso are all plausible candidates) and there is no when-not guidance, so the routing decision is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kb_checkB

Lint a field note: required sections + prohibited-content scan (no binary dumps, no media links). Run before opening a PR.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses what the lint checks for (required sections, binary dumps, media links), which is useful. However, it does not state whether lint failures block the PR, whether it mutates state, what permissions are needed, or what severity levels exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with a specific instruction at the end. It is appropriately sized—no filler, no repetition. Minor room for improvement in separating the checks into a clearer list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required. However, for a lint tool with no annotations and an undocumented parameter, the description is incomplete: it omits what constitutes a failure, how the note is supplied (file path vs. raw text), and whether the tool is read-only or mutates the note. These gaps matter for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, "note," with 0% schema description coverage and no description in the input schema. The tool description does not explain the expected format (inline text vs. file path), length limits, or content expectations. With coverage this low, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ("Lint") and resource ("field note"), and names the exact checks performed (required sections + prohibited-content scan). It is distinguishable from siblings like kb_new, kb_search, and kb_index, though it does not explicitly compare itself to them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one explicit timing cue: "Run before opening a PR." That establishes a workflow context but does not explain when not to use it, nor does it name alternative sibling tools (e.g., kb_new, kb_index) that might be more appropriate for other note-related tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kb_indexA

Regenerate knowledge/INDEX.md from all notes' front matter.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It usefully discloses that INDEX.md is rebuilt from front matter (an overwrite-style regeneration), but says nothing about whether the existing file is clobbered, error conditions, or any prerequisite state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the action first and the data source second. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and no parameters exist to document. For a simple regeneration tool the description is nearly complete, marred only by the absence of any when-to-use or overwrite caveat.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing further the description could add on parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (regenerate) and resource (knowledge/INDEX.md) plus the source of truth (notes' front matter). It reads differently from sibling kb_search/kb_new/kb_check, though it doesn't explicitly name a sibling to distinguish itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no alternatives named, even though kb_check and kb_search exist in the same knowledge-base family and could plausibly overlap. Usage must be inferred entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kb_newB

Create a field-note template for this rebuild. Fill it while working, then kb_check before sharing. One note per os+component milestone.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
agentNoai-agent
titleYes
statusNoworking
os_nameYes
componentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it is thin. It never states that this persists/creates durable state, whether an existing note for the same os+component is overwritten or errors, or what permissions are needed. 'Fill it while working' hints at a persistent artifact but discloses nothing concrete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and zero filler. It is well-sized, though the second sentence crams two distinct ideas (fill-then-check workflow and the one-note granularity rule) into one line.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a 6-parameter creation tool with zero annotation coverage and 0% parameter documentation, the description omits too much. Nothing is said about persistence, idempotency, or the meaning of the majority of its inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 6 parameters, so the description must compensate and largely does not. 'os+component milestone' indirectly hints at os_name and component, but tags, agent, status, and title are entirely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a field-note template') with a scope qualifier ('for this rebuild', 'One note per os+component milestone'). It distinguishes itself from the read-oriented kb siblings by naming kb_check as the follow-up step, though it never explicitly contrasts itself with kb_search/kb_path/kb_index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit workflow: 'Fill it while working, then kb_check before sharing,' and constrains granularity to 'one note per os+component milestone.' This routes the agent to a sibling tool and bounds usage, though it offers no explicit when-not-to-use or alternative creation paths.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kb_pathA

Report which knowledge-base roots resolve and where new notes go (repo checkout / packaged seed / UOS_HOME). Use when kb_search comes back empty and you need to know why.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the diagnostic nature and the resolution sources it inspects, implying a read-only probe, but never states explicitly that it has no side effects or that it requires no auth/setup. For a zero-parameter read tool the risk is low, so this is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The what (which roots resolve, where notes go) is front-loaded ahead of the when-to-use condition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is unnecessary, and the description supplies the purpose and the triggering scenario. Given the trivial input surface, little else is needed; only an explicit read-only/no-side-effects statement would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description correctly introduces no parameter concepts and instead spends its words on what gets reported.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Report) and a precise resource: which knowledge-base roots resolve and where new notes land. The parenthetical (repo checkout / packaged seed / UOS_HOME) makes the output scope concrete and distinguishes it from siblings like kb_search or kb_index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger condition — 'Use when kb_search comes back empty and you need to know why' — which routes the agent correctly and names the related tool. It lacks a when-not clause or mention of kb_check/kb_index as alternatives, so it stops short of a full routing guide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

regression_logB

Append a dated entry to tests/REGRESSIONS.md (status: pass|regressed|fixed|wip).

ParametersJSON Schema
NameRequiredDescriptionDefault
noteYes
statusYes
workspaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the key behavioral trait that this is an additive "append" to a fixed file, but says nothing about whether the file/workspace must pre-exist, permissions, or overwrite risk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence front-loads the action and target file, with the enum parenthetical earning its place. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, but with no annotations and 0% schema coverage the description leaves workspace scoping and note semantics undocumented. Adequate but with clear gaps for a three-parameter write tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with three required params, so the description must compensate. It usefully enumerates the status values (pass|regressed|fixed|wip), which the schema does not, but leaves both "workspace" and "note" entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a specific verb ("Append") with a precise resource ("tests/REGRESSIONS.md"), making the action unambiguous. It stops short of differentiating itself from sibling tools like kb_new or test_app_in_vm, which an agent may conflate with log-writing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as kb_new or compat_report, nor any prerequisite context. Usage is only inferable from the tool name and target file.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scaffold_componentB

Scaffold a ReactOS-style component: CMakeLists.txt, .spec export table, .c with DllMain, README, and an API test file.

exports: comma-separated export names from the API surface. Implementation always comes from behavioral specs written in spec/ — never from decompiled bodies.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNowin32dll
nameYes
exportsNo
workspaceYes
behavior_noteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses one meaningful trait (implementation is sourced from behavioral specs, never decompiled bodies), but omits whether scaffolding overwrites existing files, whether it is idempotent, or what prerequisites the workspace must meet. Useful but incomplete for a file-generating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and artifact list, followed by a compact note on the exports parameter and the sourcing rule. Every sentence earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that generates a multi-file component with 5 parameters and 0% schema coverage, the description leaves most parameters (kind, workspace, behavior_note) undefined and says nothing about preconditions or overwrite semantics. The output schema covers return values, but the input side remains substantially underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate. It explains only 'exports' ('comma-separated export names from the API surface') and leaves kind, name, workspace, and behavior_note entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Scaffold') plus a specific resource ('ReactOS-style component') and then enumerates the exact artifacts produced (CMakeLists.txt, .spec export table, .c with DllMain, README, API test file). That artifact list is distinctive enough to separate it from siblings like gen_api_stub or gen_compat_layer without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the closely related siblings (gen_api_stub, gen_compat_layer, build_plan). The one directive present — that implementation comes from spec/ and never from decompiled bodies — is a provenance constraint on output, not a when-to-use/alternative-selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test_app_in_vmB

Full app-compat cycle on the rebuilt (or reference) OS: copy the app in, run install_cmd then the app, capture exit/output, screenshot, record PASS/FAIL.

install_cmd/run_cmd: single command lines (argv[0] + args). Results go to tests/results/ and tests/RESULTS.md in the workspace. Use reference apps the user owns (RULES.md).

ParametersJSON Schema
NameRequiredDescriptionDefault
vmYes
argsNo
labelNo
run_cmdNo
timeoutNo
app_pathNo
workspaceNo
install_cmdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it does disclose the concrete side effects (copies the app into a VM, runs an installer and the app, captures exit/output, screenshots) and where artifacts land (tests/results/, tests/RESULTS.md). However, it says nothing about VM state mutation, cleanup, or error/timeout handling for what is clearly a side-effecting orchestration tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the workflow, then two focused second-paragraph clauses on command format and artifact locations. Reasonably tight with little filler, though the arg-format note is terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is unnecessary, and the description does convey the overall flow and artifact destinations. But for an 8-parameter, annotation-free, side-effecting tool it leaves most parameter meanings and state-mutation semantics unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters, so the description must compensate; it only clarifies install_cmd and run_cmd as 'single command lines (argv[0] + args)'. The other six parameters (vm, args, label, timeout, app_path, workspace) get no explanation anywhere, leaving most of the input surface undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific composite workflow ('full app-compat cycle') and enumerates its steps (copy app, run install_cmd, run app, capture exit/output, screenshot, PASS/FAIL). That clearly distinguishes it from a single-command sibling like vm_exec. It stops short of explicitly naming which sibling to prefer when a full cycle is not needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the workflow framing (it is the end-to-end compat test) and by the pointer to RULES.md for reference apps. There is no explicit when-to-use-this-vs-vm_exec or when-not-to guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_syscallsA

Write a concrete NT-syscall/API tracing plan for this VM (ETW/WPR steps, KDNET setup, static-surface fallback) and optionally kick off in-guest prep via qemu-ga.

Dynamic tracing needs in-guest tooling; this tool gives the exact commands and can fetch the results with vm_get_file afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
vmNo
workspaceNo
target_appNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the tool optionally performs a side-effecting action ('kick off in-guest prep via qemu-ga') and that it is primarily a plan generator, but it does not describe what that prep entails, permission/auth needs, or whether it mutates the VM state. Partial behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary action and deliverable, with the follow-up workflow noted second. The parenthetical is dense but earns its place by naming the concrete tracing mechanisms. Minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format needn't be explained. However, for a tool with an optional side-effecting in-guest prep step and no annotations, the description leaves the prep's effects and all three parameters unexplained, so it is only marginally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters, so the description must compensate. It only obliquely implies the 'vm' parameter ('for this VM') and gestures at target_app via 'NT-syscall/API tracing', leaving 'workspace' entirely unaddressed. Insufficient for full compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Write a concrete NT-syscall/API tracing plan for this VM') and enumerates the concrete deliverables (ETW/WPR steps, KDNET setup, static-surface fallback). The second sentence distinguishes it from execution siblings by clarifying it produces commands rather than performing tracing directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: dynamic tracing needs in-guest tooling, so use this to get exact commands, and it names the follow-up sibling (vm_get_file) for retrieving results. It does not state explicit when-not conditions or contrast with static alternatives like dump_api_surface, but the workflow guidance is solid.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_bootA

Boot the VM headless (boot='d' = CD-ROM first, 'c' = disk). Waits, then reports status.

Take a screenshot afterwards to see the display. Under pure TCG (no /dev/kvm) a full Windows install can take 30-90+ minutes — snapshot right after setup milestones.

ParametersJSON Schema
NameRequiredDescriptionDefault
bootNod
nameYes
wait_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden and does real work: it discloses that the VM runs headless, that the call blocks and waits before reporting status, and that a full Windows install under pure TCG can take 30-90+ minutes. It omits any statement about required permissions or whether the boot state change is reversible, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action and the boot-code mapping are front-loaded in the first sentence, and the second paragraph is short. The screenshot and TCG timing advice is useful but reads as slightly incidental rather than tightly integrated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers the hard parts: boot-code semantics and realistic timing expectations. The gap is the unexplained wait_seconds parameter and the absence of a stated prerequisite that the VM must already be created and running.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is no enum on 'boot', so the description's decoding of 'd' = CD-ROM first and 'c' = disk is the only place that meaning exists and is genuinely load-bearing. However, 'name' and 'wait_seconds' (the default 45s wait) are left entirely undocumented in both schema and description, so only one of three parameters is explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Boot the VM headless') and immediately disambiguates the cryptic 'boot' values, which separates it cleanly from siblings like vm_create, vm_status and vm_restore. An agent knows exactly what this call does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete follow-up routing: 'Take a screenshot afterwards' points to vm_screenshot and 'snapshot right after setup milestones' points to vm_snapshot. It does not state prerequisites (VM must already exist) or when not to use it, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_createB

Create a registered sandbox VM (qcow2 disk + config) for the reference ISO.

Networking is OFF by default — enabling it needs the user's OK (RULES.md). Call vm_boot next. workspace: path of the universal-os workspace this VM belongs to.

ParametersJSON Schema
NameRequiredDescriptionDefault
cpusNo
nameYes
forceNo
ram_mbNo
disk_gbNo
networkNo
iso_pathYes
workspaceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It does disclose two meaningful traits — networking defaults to off and enabling it requires user approval per RULES.md — but says nothing about what 'force' does, permission requirements, or what 'registered' implies in the system.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines, front-loaded with the core action before the networking caveat and the next-step pointer. Efficient, though the trailing 'workspace:' clause reads as a loose dangling note rather than a structured parameter statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for an 8-parameter creation tool with zero annotation coverage and 0% schema description coverage, the description leaves most parameter behavior and mutation semantics undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters, so the description must compensate and largely does not. It clarifies only 'workspace' (the universal-os workspace the VM belongs to) and indirectly 'network'; 'force', 'cpus', 'ram_mb', 'disk_gb', and 'name' semantics are left entirely to the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a registered sandbox VM') plus the artifacts produced (qcow2 disk + config) and the target (reference ISO). This distinguishes it from vm_boot, vm_exec, and other lifecycle siblings, though it doesn't explicitly contrast with snapshot/restore paths for pre-existing VMs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a sequencing hint ('Call vm_boot next') and a policy constraint ('enabling it needs the user's OK'), which is real routing value. However, it never says when NOT to create (e.g. existing VM → vm_status) or names an alternative for cloning/restoring instead of fresh creation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_execB

Run a command inside the guest via the QEMU guest agent (qemu-ga must be installed).

args: single string, split on whitespace, quotes respected, backslashes kept as-is (Windows guest paths like 'C:\uos-tests\app.exe' survive). Windows guests: use 'cmd.exe' with args like '/c dir C:'. If the agent is missing the error explains how to install it (virtio-win ISO for Windows guests).

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
nameYes
pathYes
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses argument-parsing behavior, backslash preservation, and that a missing agent produces an explanatory error, but it does not describe the execution profile (arbitrary code execution, permission requirements) or what the timeout parameter controls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action before args mechanics and Windows-specific notes; each sentence carries information. Slightly dense in the middle but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the agent-requirement plus arg semantics are covered. However, with no annotations and 0% schema coverage, the timeout and name/path semantics remain unexplained, leaving gaps for an execution tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for all four parameters. It explains the 'args' parameter thoroughly (whitespace splitting, quote handling, Windows paths) but leaves 'name', 'path', and 'timeout' completely undocumented, so only one of four parameters is addressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Run a command inside the guest') and the mechanism (QEMU guest agent), which lets an agent distinguish it from siblings like vm_sendkey or test_app_in_vm. It does not explicitly name a sibling it is not, keeping it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides usage context (qemu-ga must be installed, Windows guests should call cmd.exe with /c) and recovery guidance when the agent is missing, but never says when to choose this tool over alternatives such as vm_sendkey or test_app_in_vm. Usage is implied rather than scoped.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_get_fileC

Pull a file out of the guest (e.g. an .etl trace, a crash dump, a test log).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
guest_pathYes
local_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and falls well short. It does not disclose that the VM must be running, that the operation writes to the host filesystem (local_path), whether existing files are overwritten, or any permission/authentication requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words; the action is stated before the examples. It is appropriately sized, though its brevity reflects under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, but with zero annotations, 0% parameter coverage, and three required params, the description leaves critical gaps: prerequisites, host-side side effects, and parameter meanings are all absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three required parameters (name, guest_path, local_path) have 0% schema description coverage, so the description must compensate and largely does not. 'Pull a file out of the guest' weakly implies guest_path is the source and local_path the destination, but 'name' is entirely unexplained and no path format is given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Pull') and resource ('a file') with a clear source/destination relationship ('out of the guest'), and the examples (.etl trace, crash dump, test log) concretely illustrate what kind of files. It does not explicitly name the inverse sibling vm_put_file, so differentiation is only implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The examples hint at a post-test/post-run retrieval scenario, but there is no explicit statement of when to use this versus vm_put_file, vm_exec, or vm_log, and no prerequisites or exclusions are given. The agent must infer the usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_logB

Tail the VM's serial console log and show the qemu command line (debugging boot).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
tailNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool reads a log and additionally surfaces the qemu command line, which hints at read-only behavior, but it says nothing about permissions, whether the log stream blocks or terminates, or how much output is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the primary action front-loaded and the debugging motivation in a short parenthetical. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But with two parameters at 0% schema coverage and no annotations, the definition leaves the agent guessing about the meaning of 'name' and the line-count semantics of 'tail', which is a real gap for an otherwise simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with two parameters ('name' and 'tail') both undocumented in the schema. The description mentions 'Tail' but only as the verb, never explaining that the integer parameter controls how many lines to return, and never clarifies that 'name' identifies the VM. It does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Tail the VM's serial console log' plus a secondary output ('show the qemu command line'). This is clearly distinguishable from siblings like vm_status, vm_screenshot, and vm_boot. It stops short of naming an alternative tool, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(debugging boot)' implies the intended scenario, giving implied usage context. However, it names no alternative (e.g. vm_status for non-log state) and states no exclusions or prerequisites, so guidance is only inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_mount_isoA

Swap the CD-ROM medium in the running VM (e.g. insert virtio-win to install the guest agent, then swap back to the reference ISO).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
iso_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It discloses that the operation happens on a running VM and is a swap (implying reversibility), but it omits permissions needed, side effects on existing medium, and error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One well-structured sentence with the purpose front-loaded and a parenthetical example that adds context without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and 0% schema coverage, the description should explain the parameters and behavioral constraints, but it only gives a use-case example. The output schema exists, so return values are covered, but invocation details remain incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither parameter is described. The mention of 'virtio-win' and 'reference ISO' hints that iso_path is an ISO file path, but 'name' (presumably the VM name) is not addressed at all; the description does not compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Swap') and resource ('CD-ROM medium in the running VM'), and the example clarifies intended media. It is distinguishable from sibling ISO inspection/verification tools because it modifies VM state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a concrete scenario ('insert virtio-win to install the guest agent, then swap back to the reference ISO'), giving clear context for when to use it. However, it does not name alternative tools or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_network_configA

Set the VM's network flag: OFF by default; enabling needs the user's OK (RULES.md). Only user-mode NAT is implemented (bridge raises honestly). Applies at the next vm_boot - a running VM keeps its old network until shutdown + boot. Check with vm_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
natNo
nameYes
enableNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description must carry the behavioral burden. It discloses the default (OFF), the authorization requirement (user's OK), the implementation limitation (NAT only, bridge raises), the timing semantics (deferred to next boot), and the verification path (vm_status). This is substantial context for a 3-param mutation tool. Minor gap: no explicit mention of reversibility or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and default, then implementation limits, then timing/verification. Efficient, though the parenthetical about bridge is slightly tangential.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need not be explained. The description covers timing, authorization, and implementation limits, but omits parameter semantics (critical given 0% schema coverage) and any note on what happens if 'enable' is true without 'nat' or vice versa.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It mentions the 'network flag' and NAT but does not explain the three parameters (nat, enable, name) or their defaults/interactions. An agent cannot infer that 'enable' is the primary toggle and 'nat' selects network mode from the description alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (set the VM's network flag) and clarifies scope limitations (only NAT implemented, bridge unsupported). It distinguishes itself from runtime-control siblings like vm_boot by clarifying the config-vs-runtime relationship, though it doesn't name explicit alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly defines when the change takes effect (next vm_boot, not a running VM) and points to vm_status for verification. It states a prerequisite (user's OK per RULES.md). Missing explicit pointers to alternative network tools, but usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_put_fileA

Copy a file from the host into the guest (via qemu-ga). Guest paths are Windows-style for Windows guests (e.g. C:\uos-tests\app.exe). Size cap: ~2GB practical.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
guest_pathYes
local_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose real behavioral traits: the transport mechanism (qemu-ga), Windows-style guest path convention, and a ~2GB practical size cap. However, it says nothing about overwrite behavior for an existing guest file, required permissions/agent availability, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core purpose, with the path convention and size cap as supporting detail. Nothing is padded, though the size-cap note could be tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers mechanism, path convention and a size limit, but for a mutation tool with zero annotation coverage it omits overwrite semantics, permission/agent requirements, and error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all three parameters. It clarifies the guest_path format (Windows-style, with an example) and implies local_path is the host-side source, but the required 'name' parameter is never explained at all, leaving a full third of the parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (copy), resource (a file), and direction (host into guest), which is exactly what separates it from the sibling vm_get_file. The qemu-ga mechanism further pins down what is happening. An agent can select this tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the host-to-guest direction tells the agent to use this when pushing files in, but no alternatives (vm_get_file, vm_mount_iso, init_workspace) are named and no prerequisites or exclusions are given. It gives path-format guidance but not when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_restoreC

Restore the running VM to a previously saved snapshot (revert experiments safely).

ParametersJSON Schema
NameRequiredDescriptionDefault
tagYes
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Restore' implies discarding current VM state, but the description never states that the current state is lost, whether the VM must be stopped first, whether permissions are required, or how long the operation takes — the '(revert experiments safely)' phrase gestures at destructiveness without disclosing it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler; the core action leads and the clarifying parenthetical follows. It is efficiently sized, though the parenthetical is doing more atmospheric than informational work.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, but for a state-mutating VM operation with no annotations and fully undocumented parameters the description is too thin. An agent cannot tell whether a snapshot must pre-exist, whether the VM is left running afterward, or how to identify the correct VM.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters (name, tag) have 0% schema description coverage, and the description adds no meaning for either — it never explains what 'name' identifies or what the 'tag' selects. With two undocumented required parameters, the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restore) and resource (running VM to a previously saved snapshot), which an agent can distinguish from siblings like vm_snapshot or vm_status. It does not, however, explicitly contrast itself with the snapshot-creation tool or clarify the restore-vs-snapshot relationship.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(revert experiments safely)' implies the intended context — undoing experimental changes — but gives no explicit when-to-use/when-not, no prerequisite that a snapshot must already exist, and no pointer to vm_snapshot as the counterpart operation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_screenshotA

Capture the VM display to a PNG and return its path.

Use it after boot, after each setup screen, and after app tests — it is the AI's eyes on the guest. Read the PNG with vision to decide the next step.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the output medium (PNG file path) and the intended downstream action (read with vision), but says nothing about preconditions — e.g. whether the VM must be powered on or a display attached — or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action, then the when, then the purpose. No filler; every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be spelled out, and the usage guidance is strong for a one-parameter tool. The remaining gap is the unexplained 'name' parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter 'name' with 0% schema description coverage, so the description must compensate and does not. It never mentions what 'name' identifies or how it relates to a VM instance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Capture the VM display to a PNG and return its path.' No sibling (vm_status, vm_log, vm_exec, etc.) captures the display, so the agent can distinguish it immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit timing triggers — 'after boot, after each setup screen, and after app tests' — which is far more than most definitions offer. It doesn't name alternatives or a when-not condition, but for a capture primitive the positive triggers are clear guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_sendkeyA

Send a key combo to the VM display: 'ret', 'esc', 'f8', 'ctrl-alt-delete', 'shift-f10', 'spc', 'tab', 'up'/'down'/'left'/'right', single letters/digits.

This drives text-mode installers BEFORE the guest agent exists (the part vm_exec cannot reach). One combo per call; screenshot after each to see the installer's reaction.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysYes
nameYes
hold_msNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses the pre-agent timing window, the one-combo-per-call constraint, and the recommended screenshot verification loop. It does not, however, mention error behavior (e.g., what happens if keys are unrecognized or the display is not focused), leaving some gaps for a mutation-style input tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the accepted key vocabulary followed by the workflow constraint and the verification tip. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers purpose, timing, and invocation cadence. The main omission is the meaning of 'name' and 'hold_ms', which leaves the call specification slightly incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does partially: the enumerated key names ('ret', 'esc', 'ctrl-alt-delete', 'shift-f10', etc.) effectively document the accepted values for 'keys'. But 'name' (presumably the VM identifier) and 'hold_ms' (default 100) are never explained, leaving two of three parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (send) and resource (a key combo to the VM display), and explicitly differentiates itself from the closest sibling by noting it drives text-mode installers 'BEFORE the guest agent exists (the part vm_exec cannot reach).' An agent can select between vm_sendkey and vm_exec without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (pre-guest-agent text-mode installers), names the alternative it is not (vm_exec), and prescribes the operating rhythm: 'One combo per call; screenshot after each to see the installer's reaction.' Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_shutdownA

Shut the VM down (ACPI powerdown; falls back to quit after a timeout; force=true kills immediately). Registry metadata is kept — vm_boot can start it again.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
forceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and mostly succeeds: it discloses the ACPI powerdown mechanism, the fallback-to-quit timeout behavior, the kill semantics of force=true, and the important side effect that registry metadata persists. It omits prerequisites such as whether the VM must be running and what happens if it is already stopped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with zero filler; the action and its mechanism lead, and the state-persistence caveat follows. Every clause carries distinct operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not needed, and the description covers mechanism, fallback, and post-shutdown state. Remaining gaps (error conditions, required VM state, permissions) are minor for a two-parameter lifecycle tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It meaningfully explains the non-obvious 'force' parameter ('kills immediately') and its place in the fallback sequence, while 'name' is self-evident from context. One parameter's semantics are effectively undocumented but the ambiguous one is well covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Shut the VM down') and disambiguates the action space by naming the inverse sibling ('vm_boot can start it again'). An agent can immediately distinguish this from vm_boot, vm_snapshot, or vm_restore without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the force=true note hints at when to escalate to an immediate kill, but there is no explicit when-to-use or when-not-to-use guidance (e.g., graceful shutdown before teardown vs. aborting a hung VM). It names vm_boot only as a recovery path, not as an alternative to weigh.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_snapshotA

Save an internal qcow2 snapshot of the running VM (e.g. 'clean-install', 'pre-app-test').

Snapshot before risky operations: installing apps, changing components, registry edits. Restore with vm_restore. Tag naming: kebab-case, one milestone per tag.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagYes
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the snapshot is 'internal qcow2' and tied to the running VM, implying a non-destructive save rather than a state mutation, and it pairs with vm_restore. It does not say whether the VM must be running, whether snapshots consume disk, whether existing snapshots are overwritten, or what the call returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines, front-loaded with the core action, then usage, then the paired tool and naming rule. No filler sentences and no repetition of the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and purpose/usage/naming are covered. However, for a tool whose two required string parameters are undocumented in the schema, the description should have resolved what 'name' means relative to 'tag' and whether snapshotting requires a specific VM state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither 'name' nor 'tag' is described in the schema. The description compensates for only one of the two: it defines the tag convention (kebab-case, one milestone per tag) and gives examples, but leaves 'name' entirely unexplained, so the crucial name-vs-tag distinction remains unresolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Save an internal qcow2 snapshot of the running VM'), names the artifact type, and explicitly distinguishes itself from vm_restore, which is the natural counterpart an agent might confuse it with. No schema inspection is needed to know what this does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use trigger ('before risky operations: installing apps, changing components, registry edits') and points to vm_restore for the reverse operation. It stops short of stating when NOT to snapshot (e.g. disk-space limits, VM must be running) but the operational context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vm_statusA

Report one VM's status, or every registered VM when name is empty.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the important behavioral trait that an empty name returns all registered VMs, but says nothing about error behavior for an unknown name, required permissions, or whether the VM must be running.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler that encodes both the primary action and the parameter-driven mode switch.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and for a one-parameter read tool the mode switch is covered. Missing only edge-case behavior (unknown name, permissions), which is minor at this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate; it does explain the key semantic that an empty/default name means 'all VMs' rather than a lookup failure. It stops short of clarifying whether 'name' is a registered name, ID, or label.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and resource ('VM's status'), plus the exact scope switch (one VM vs every registered VM). No sibling in the list offers VM status reporting, so an agent can route here without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence implicitly defines two invocation modes (named VM vs all VMs), which is useful context, but it names no alternatives, prerequisites, or when-not-to-use conditions. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_statusA

Summarize the rebuild workspace: required dirs (src/spec/api/tests/reports/ third-party), the WORKSPACE.md journal state and its last line, api/ surface count and src/ components. Use to decide the next phase of the loop.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose concretely what the tool inspects (directory presence, journal state and last line, API surface count, component list). It stops short of stating that it is a non-mutating read or any permission/side-effect behavior, but the enumerated inspection targets are unusually informative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the concrete summary contents before the usage note, with no filler. Efficient, though the enumeration of dirs is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values are already covered, and the low-complexity single-param, non-nested schema keeps the burden light. The main remaining gap is the undocumented workspace parameter, but the description otherwise fully explains the tool's scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter "workspace" has 0% schema description coverage, and the description never explains what value it expects (path, name, identifier) or its format. The phrase "the rebuild workspace" hints at the target but adds no actionable semantics beyond the parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ("Summarize") and resource ("the rebuild workspace") and enumerates exactly what it reports: required dirs, WORKSPACE.md journal state and last line, api/ surface count, and src/ components. It is clearly distinguishable from init_workspace/cleanup_workspace by its read-and-report nature, though it never names a sibling to contrast against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use to decide the next phase of the loop" gives an implied usage context, but there is no explicit when-not guidance and no routing to alternatives such as build_plan or compat_report. An agent gets a rough sense of when to reach for it but must infer the boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 39 tool updatesv0.2.1
    • First observedanalyze_pe
    • First observedbehavior_diff
    • First observedbuild_plan
    • First observedcleanup_workspace
    • First observedclone_reference_repo
    • First observedcompat_report
    • First observeddump_api_surface
    • First observedenv_check
    • First observedfingerprint_windows
    • First observedgen_api_stub
    • First observedgen_compat_layer
    • First observedinit_workspace
    • First observediso_extract
    • First observediso_inspect
    • First observediso_verify
    • First observedkb_check
    • First observedkb_index
    • First observedkb_new
    • First observedkb_path
    • First observedkb_search
    • First observedregression_log
    • First observedscaffold_component
    • First observedtest_app_in_vm
    • First observedtrace_syscalls
    • First observedvm_boot
    • First observedvm_create
    • First observedvm_exec
    • First observedvm_get_file
    • First observedvm_log
    • First observedvm_mount_iso
    • First observedvm_network_config
    • First observedvm_put_file
    • First observedvm_restore
    • First observedvm_screenshot
    • First observedvm_sendkey
    • First observedvm_shutdown
    • First observedvm_snapshot
    • First observedvm_status
    • First observedworkspace_status

TDQS

B3.4/5.0

Scored across 39 tools

Disambiguation4/5

Most tools have clearly distinct purposes (vm_snapshot vs vm_restore, kb_search vs kb_path, build_plan vs env_check). There is mild overlap in the ISO cluster (iso_inspect/iso_verify/iso_extract/fingerprint_windows) and between env_check/workspace_status/build_plan which all report environment state, but descriptions make the boundaries readable.

Naming Consistency4/5

Predominantly consistent verb_noun snake_case with clear namespace prefixes (vm_, iso_, kb_). A few deviations exist: fingerprint_windows is noun_verb, and behavior_diff/compat_report/regression_log are noun-only, but all remain readable and predictable.

Tool Count3/5

39 tools is heavy and above the comfortable range, though the domain (VM management, ISO/binary analysis, scaffolding, testing, knowledge base) is genuinely broad and most tools map to a real workflow step. It is on the borderline of over-fragmentation rather than gratuitous sprawl.

Completeness4/5

Covers the full lifecycle: VM create/boot/snapshot/restore/shutdown, ISO inspection, API-surface analysis and diffing, workspace scaffolding, app-compat testing, and KB management. Minor gaps (e.g. no VM delete/deregister, no in-place file edit) are workable around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    A lightweight server that enables AI agents to interact with the Windows operating system, allowing for file navigation, application control, UI interaction, and QA testing through various tools.
    7,620
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A local sandbox that provides AI agents with code execution, filesystem access, and a full GUI desktop environment for 'computer use' capabilities. It exposes tools for shell commands, multi-language code interpretation, and remote desktop automation via the Model Context Protocol.
    212
    GPL 3.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI clients to securely control and interact with a local Windows machine through 218 configurable tools for files, Git, processes, Windows UI, browser automation, WSL, Office, recovery, skills, and child MCP servers.
    11 npm
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables local AI models to perform defensive cybersecurity analysis through narrowly scoped read-only tools for host posture, Windows security operations, file/IOC triage, code scanning, and allowlisted filesystem/network access while enforcing boundaries and audit trails.
    -