Skip to main content
Glama

Universal Game Modder

One MCP server to inspect and mod any game — driven by an AI agent.

Universal Game Modder (UGM) is a local Model Context Protocol server that gives an AI agent (like Claude in Claude Code) hands for game reverse-engineering and modding: it auto-detects a game's engine, then exposes a toolset for analyzing and patching the binaries — all running locally on your own machine, no cloud, no telemetry.

This is the free open-core edition (v0.1.3).


⚠️ Responsible-use notice

UGM is a general-purpose binary-analysis and interoperability tool. It is meant for modding, interoperability, security research, and education on software you legally own or are authorized to analyze. You are responsible for how you use it. Before installing, read ACCEPTABLE_USE.md and LICENSE-EULA.md.


Related MCP server: memscope-mcp

What's in the free edition

  • One stdio MCP server that an agent connects to and drives.

  • Game-engine auto-detection — point it at a game directory; it identifies Unity (Mono / IL2CPP), Unreal, Godot, Java, and native binaries from file signatures.

  • Native binary analysis + patching toolset — PE header analysis, hex read/write/search/replace, IDA-style pattern scanning, string extraction, checksums, and binary diffing, running natively in-process (no external tools required).

The guided game-modder workflow skill and the delegated decompilation backends land in a later release; this edition is the MCP server + engine detection + the native toolset.

Not in the free edition (Pro tier)

The web dashboard, the delegated Unity/Unreal/JAR decompilation backends, and native disassembly are part of the Pro tier and are not included here. The free edition is fully functional on its own for engine detection and native binary work.


Requirements

  • Node.js 18+ (uses ES modules and better-sqlite3)

  • Claude Code or any MCP-capable client

  • Windows / macOS / Linux (native tools are cross-platform; the smallest cut ships no OS-specific binaries)


Install

# from the repo root
./setup.ps1

Or manually (dist/ ships prebuilt; there is nothing to compile in this edition):

npm install

Then register the server with your MCP client. For Claude Code, add to your MCP config:

{
  "mcpServers": {
    "universal-game-modder": {
      "command": "node",
      "args": ["dist/index.js"]
    }
  }
}

On first run, ugm.config.json ships with empty placeholders — the free edition needs no external paths. (The Pro delegate backends are where those get filled in.)

Config resolution

The server reads the first of these that exists and merges it over built-in defaults:

  1. the file named by the UGM_CONFIG environment variable

  2. ugm.config.local.json next to package.json — your machine-specific paths; keep it out of version control (the repo's .gitignore already does)

  3. ugm.config.json — the tracked file, placeholders only

So you can fill in backend paths without ever editing a file that could end up in a release or a pull request.


Quickstart

Once connected, ask your agent to work through these. Every example below is executed against a real binary before each release — see Verified examples.

Parameter naming: file-level tools take file_path (and file_path_a / file_path_b for comparisons). Game-directory tools take game_path. Disassembly tools take binary_path.

1. Detect what a game is built with:

detect_engine  { "game_path": "C:\\Path\\To\\Game" }

2. Set it as the active target:

load_game  { "game_path": "C:\\Path\\To\\Game" }

3. Analyze a binary:

analyze_file_format  { "file_path": "...\\SomeBinary.dll" }   # magic bytes, managed vs native
analyze_pe_full      { "file_path": "...\\SomeBinary.exe" }   # PE headers, sections, data directories

4. Search and inspect:

extract_strings          { "file_path": "...", "min_length": 8 }
extract_dll_classes      { "file_path": "...", "search_terms": ["Health","Damage"] }
pattern_scan             { "file_path": "...", "pattern": "48 8B ?? ?? ?? ?? ??" }
search_binary_pattern    { "file_path": "...", "patterns": ["maxHealth"] }

5. Patch and verify:

hex_read       { "file_path": "...", "offset": 4096, "length": 64 }
hex_replace    { "file_path": "...", "search_hex": "90 90", "replace_hex": "EB 00" }
calculate_checksums       { "file_path": "..." }              # before/after integrity
compare_binaries_detailed { "file_path_a": "...", "file_path_b": "..." }

6. Disassemble (native, included):

disassemble_function  { "binary_path": "...\\SomeBinary.exe", "rva": 4096 }
disassemble_range     { "binary_path": "...\\SomeBinary.exe", "rva": 4096 }

Full free-edition tool list

These 29 tools run entirely in-process and need no external backend.

Native binary tools (19) — all take file_path unless noted: Detection / PE: analyze_pe_full, analyze_file_format, analyze_dll_structure, rva_to_offset (+rva), offset_to_rva (+offset) Hex: hex_read (+offset), hex_write (+offset,hex_data), hex_search (+hex_pattern), hex_replace (+replace_hex) Scanning: pattern_scan (+pattern), pattern_scan_all (+pattern), search_binary_pattern (+patterns) Strings / classes: extract_strings, extract_strings_advanced, extract_dll_classes Integrity / diff: calculate_checksums, compare_binaries_detailed (file_path_a,file_path_b) Godot: analyze_godot_pck Disassembly: disassemble_function (binary_path,rva)

Workflow / session tools (10): detect_engine, load_game, game_status, find_steam_games, mod_this_game, find_gameplay_values, build_and_deploy, debug_mod, scaffold_mod, list_available_tools

What the server also lists (Pro backends)

For transparency: the server advertises 100 tools total. The other 64 are the Unity (27), Unreal (11), and Java/JAR (26) decompilation suites, which route to external Pro-tier backend executables. They appear in the tool list, but calling one without a configured backend returns a clear error:

ERROR: Backend unity-decompiler not configured (missing executable path)

That is expected behavior in the free edition, not a defect. Configure their paths in ugm.config.json (Pro) to enable them.

Verified examples

The Quickstart calls above are not aspirational. Each release is gated on a verification pass that executes them against a real PE binary and requires every one to return real data. The v0.1.3 pass (2026-09-02, on the restored release cut) ran the gate's 20 documented calls plus 3 clean-failure checks; results are recorded in VERIFICATION.md.


License

The UGM code is released under the MIT License. Use of the tool is additionally governed by LICENSE-EULA.md and ACCEPTABLE_USE.md. Third-party components used by the Pro-tier backends are attributed in licenses/.

Changelog

See CHANGELOG.md. Current version: v0.1.3 — the config can no longer ship machine paths, and the release was restored and re-verified after a disk incident. v0.1.2 fixed engine detection; v0.1.1 fixed the docs and added the gate.

Available Tools

100 tools
analyze_bepinex_logC

Analyze BepInEx log for errors, warnings, and patch issues.

ParametersJSON Schema
NameRequiredDescriptionDefault
log_pathYesPath to BepInEx LogOutput.log

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state whether the log must exist, whether analysis is read-only, what format the findings take, or how errors versus warnings are treated. For a diagnostic tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the resource front-loaded. Nothing is wasted, though it is perhaps too terse given the absence of annotations and usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should disclose return shape or at least the nature of the analysis output. It says nothing about what a caller receives or how to act on the findings, leaving key behavior undefined for a diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is fully documented in the schema. The description adds no format details (e.g., what if the file is missing or not a valid log), so baseline 3 applies when the schema already does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Analyze) and resource (BepInEx log), and names what it extracts (errors, warnings, patch issues). However, it does not distinguish itself from siblings like verify_patches or debug_mod, which could plausibly also surface patch issues, so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. Given the large set of sibling tools, an agent has no signal for when log analysis is the right move versus, say, verify_patches or debug_mod.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_dll_structureB

Analyze the overall structure of a DLL: sections, imports, exports summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPath to DLL file

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the inspection categories (sections, imports, exports), which is useful, but says nothing about read-only safety, failure on non-DLL input, size limits, or output form. Partial behavioral coverage only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb and resource first, followed by the inspected artifacts. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description is the only disclosure channel, yet it does not describe the return shape or error behavior. The scope enumeration (sections, imports, exports) is the minimum needed to call it, but the definition is not complete for a binary-analysis tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single file_path parameter, so the schema already documents the input. The description adds no format, path-resolution, or validation detail beyond it, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Analyze the overall structure of a DLL') and enumerates what is inspected (sections, imports, exports summary), so the agent knows exactly what it produces. However, it does not distinguish itself from overlapping siblings like analyze_pe_full or extract_dll_classes, which also inspect binary structure.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no reference to alternatives such as analyze_pe_full or analyze_file_format, even though several siblings analyze binaries. The agent is left to infer when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_file_formatB

Detect file format from magic bytes. Works with PE, ELF, Mach-O, ZIP, Java class, and other formats.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPath to the file

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the detection mechanism (magic bytes) and supported formats, which is meaningful context, but says nothing about the return shape, whether it reads only a header, or how it handles unknown formats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, zero waste, with the core action front-loaded before the format enumeration. Nothing redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, non-mutating detection tool with no output schema, the description covers the essentials: what it detects, how, and which formats. Its only gap is the absence of any output-format or fallback behavior hint, which is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (file_path) at 100% schema description coverage, so the schema already documents it. The description adds no syntax, path-format, or edge-case detail beyond what the schema provides, matching the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Detect file format') and the mechanism ('from magic bytes'), plus the covered formats (PE, ELF, Mach-O, ZIP, Java class). It is clearly distinct from deeper analysis siblings like analyze_pe_full, though it doesn't name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives are given. The format list hints at applicability but the agent gets no instruction on when to reach for this versus analyze_pe_full, detect_engine, or calculate_checksums.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_godot_pckC

Analyze a Godot .pck archive and list its contents.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional filename filter
file_pathYesPath to .pck file

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether analysis is read-only, whether large archives are paginated, how contents are structured in the response, or any limits. 'Analyze' and 'list its contents' give only the vaguest behavioral signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and target. No filler, no redundant restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 2-parameter tool with no annotations and no output schema: the description should explain what 'listing contents' returns (file entries? asset paths? sizes?) and any behavioral caveats. It leaves the agent guessing about the response shape and safe usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – both file_path and filter are documented in the schema. The description adds no format details, accepted extensions, or filter syntax beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Analyze) and resource (Godot .pck archive) with the added outcome 'list its contents'. Clear enough to distinguish from generic tools like analyze_file_format or analyze_pe_full, though it doesn't name those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this versus siblings such as detect_engine, unpack_game, decode_assets, or read_asset, all of which could operate on game archives. No preconditions or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_pe_fullB

Full PE (Portable Executable) analysis. Shows headers, sections, data directories, and whether it is managed (.NET) or native.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPath to PE file (.exe or .dll)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states what the analysis shows, implying a read-only inspection, but does not explicitly confirm non-destructive behavior, required permissions, or failure modes. It adds output-category context but leaves the safety profile unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences, front-loaded with the core purpose and then the output summary. No filler or redundancy; every sentence contributes directly to understanding what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema and no annotations, the description gives a reasonable high-level preview of returned information (headers, sections, directories, managed/native). It does not cover error cases or return format, but the core analysis scope is adequately communicated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single file_path parameter is fully described in the schema as 'Path to PE file (.exe or .dll)'. The description adds no extra meaning about the parameter, such as format constraints or path resolution. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('analysis') and resource ('PE (Portable Executable)'), and enumerates what it reveals: headers, sections, data directories, managed/native. This distinguishes it from generic format analysis, but it does not explicitly name or differentiate from sibling tools such as analyze_dll_structure or analyze_file_format.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or comparison to alternatives like analyze_file_format or analyze_dll_structure. The name implies PE-specific analysis, but the description provides no conditions, prerequisites, or exclusions for selecting this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_save_formatC

Analyze game save system by examining SaveManager classes.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It reveals only the mechanism (examining SaveManager classes); it does not state that the operation is read-only/non-destructive, what permissions or loaded state it requires (the required 'assembly' key implies prior loading), or what the analysis yields. 'Analyze' is left completely unqualified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. It is appropriately sized, though the brevity overlaps with under-specification rather than pure economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations exist, so the description must explain what the analysis returns (e.g., save structure, serialization details, detected SaveManager types), and it does not. For an analysis tool that gates other workflows, this leaves the agent guessing about the result shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter exists and schema description coverage is 100% ('assembly: Assembly key'), so the schema already documents it. The description adds nothing about the parameter, but per the high-coverage rule the schema doing the work makes a baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (Analyze) and resource (game save system), plus the mechanism (SaveManager classes), so an agent knows the target. However, it does nothing to distinguish itself from adjacent analysis siblings like analyze_file_format, analyze_dll_structure, or analyze_pe_full.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative routing guidance. The relationship to load_assembly (which presumably supplies the assembly key) or to save-related siblings is left entirely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_and_deployC

Compile a mod project and deploy the output to the game mods directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
deploy_pathNoPath to deploy the built mod (auto-detected if omitted)
project_pathYesPath to the mod project directory
build_commandNoCustom build command (auto-detected if omitted)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether deployment overwrites existing files in the mods directory, what permissions are needed, whether the build is destructive, or what a failure looks like — significant gaps for a mutating build/deploy tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the two core actions. Efficient with no filler, though it is arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating build-and-deploy tool with no annotations and no output schema, the description omits when to use it, what the deployment does to existing files, and what the result looks like. It leaves too much to inference for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (project_path, deploy_path, build_command) are already documented in the schema, including their auto-detection behavior. The description adds nothing beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs (compile, deploy) and the resource (a mod project plus the game mods directory), so the agent knows roughly what it does. However, it does not differentiate from siblings like compile_plugin or scaffold_mod, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as compile_plugin or mod_this_game, and no prerequisites or ordering constraints are stated. Usage is only implied by the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calculate_checksumsC

Calculate MD5, SHA1, SHA256 checksums for a file.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPath to the file

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only calculation but does not state whether the file is modified, what permissions are needed, whether all three checksums are always returned, or how large files are handled. For a tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, and the core action and algorithms are front-loaded. It is efficient, though extremely terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter calculation tool the description covers the core action and algorithms, but with no output schema it omits what is returned (e.g., all three checksums together, format, empty-file behavior) and gives no usage context. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single parameter, so the schema already documents 'file_path' as the path to the file. The description adds no format, validation, or path-handling details beyond what the schema provides, making 3 the appropriate baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Calculate') and resource ('checksums for a file') and names the three algorithms, so an agent knows exactly what the tool produces. It does not explicitly differentiate itself from siblings, but no sibling computes checksums, so the purpose is still clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no alternatives. It is a bare statement of capability, leaving usage entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_binaries_detailedC

Compare two binary files and show differences.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_diffsNoMax differences to show (default 50)
file_path_aYesPath to first file
file_path_bYesPath to second file

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read-only comparison but does not state side effects, permissions, input constraints, or output format, leaving meaningful behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is concise, though its brevity is a symptom of under-specification rather than optimal efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a binary-diff tool with no output schema and no annotations, the description is too thin. It does not explain what 'differences' are shown, how detailed the output is despite the name, or how it relates to sibling comparison tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both required file paths and the optional max_diffs parameter. The description adds no parameter meaning beyond the schema, making the baseline score of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Compare') and resource ('two binary files') and says it shows differences. It is clear but does not differentiate itself from siblings like diff_assemblies, jar_diff, or compare_signatures, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives. It does not mention use cases, prerequisites, or exclusions, leaving the agent to infer context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_signaturesC

Compare expected Harmony patch signature with actual target method.

ParametersJSON Schema
NameRequiredDescriptionDefault
patch_kindNoPatch kind: prefix or postfix (default: prefix)
target_typeYesTarget type name
patch_paramsYesPatch parameter types (comma-separated)
game_assemblyYesAssembly key
target_methodYesTarget method name

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing beyond the one-line purpose. It does not say whether this is a read-only analysis, whether it mutates the assembly, what permissions or loaded-assembly state are required, or what a comparison result looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, which is appropriate for a narrowly scoped tool. It is terse rather than padded, though it borders on under-specified for a 5-parameter operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description is the only place behavioral and result context could appear, and it is silent on both. For a patch-signature comparison tool the agent gets no indication of what a result means (match/mismatch/diagnostic) or what state the target assembly must be in.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each of the five parameters documented (patch_kind, target_type, patch_params, game_assembly, target_method), so the baseline of 3 applies. The description adds no syntax, format, or default information beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (compare) and the two resources being compared (expected Harmony patch signature vs actual target method), so an agent knows roughly what the tool does. However, it does not differentiate itself from close siblings like validate_patch_target or verify_patches, which likely overlap in the patch-validation space.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as validate_patch_target or verify_patches. The agent is left to infer when signature comparison is the right call versus other patch-validation tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compile_pluginC

Compile a BepInEx plugin C# source file using Roslyn.

ParametersJSON Schema
NameRequiredDescriptionDefault
output_pathYesOutput DLL path
source_filesYesC# source file path(s), semicolon-separated
assembly_nameNoAssembly name (optional)
managed_directoryYesPath to game Managed directory
bepinex_core_directoryYesPath to BepInEx/core directory

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not say what happens on compilation failure (are Roslyn diagnostics returned?), whether output_path is overwritten, whether reference resolution failures are fatal, or what the response looks like. For a compile step this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler and the core action front-loaded. It is not padded, but it is also thin enough that the brevity comes partly from omission rather than efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters, 4 required, no annotations, and no output schema, the description should explain failure behavior and the shape of a successful compile result. It leaves the agent without enough context to anticipate errors or interpret success.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema. The description adds only the Roslyn implementation detail and no additional semantics about parameter interaction (e.g., how source_files relates to managed_directory references). Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: compiling a BepInEx plugin C# source file, and names the compiler (Roslyn). It is distinguishable from siblings like generate_plugin and build_and_deploy, though it does not explicitly say how it differs from them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus generate_plugin, build_and_deploy, or validate_assembly, nor any stated prerequisites. The agent must infer that this is the raw compile step from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debug_modC

Analyze mod logs to diagnose errors and suggest fixes.

ParametersJSON Schema
NameRequiredDescriptionDefault
log_pathNoPath to the mod log file
game_pathNoGame directory (for auto-detecting log location)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It doesn't state that it only reads files, what permissions/paths it needs, how the two optional params interact, or what output form the 'fixes' take. For a diagnostic tool with zero annotation coverage this is a real gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the purpose front-loaded and nothing wasted. It is efficient, though its brevity contributes to the missing behavioral detail rather than compensating for it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description is the only source of behavioral context, yet it omits what the analysis returns, whether fixes are applied or merely suggested, and how the optional params interact. Inadequate for a 2-param diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (log_path, game_path) are already documented in the schema; baseline 3 applies. The description adds no meaning beyond it, and notably does not clarify how the two optional params relate (log_path for explicit logs vs game_path for auto-detection).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: "Analyze mod logs to diagnose errors and suggest fixes" states what it does and its output intent. It does not distinguish itself from the close sibling analyze_bepinex_log, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives. Given the sibling analyze_bepinex_log, the agent gets no signal about which log-analysis tool to pick or when this one applies instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decode_assetsA

Universal Asset Decoder. Reads a PR-1 autopsy catalog, cracks open Unity .assets containers, and decodes their members back into The Model (decoded=1, decoded_path set): Texture2D -> PNG (streamed pixels from sibling .resS), Mesh -> glTF .glb + .obj (plain/uncompressed; compressed noted+deferred), and AudioClip -> WAV (PCM) / OGG (bare Vorbis); FSB5-wrapped Vorbis + other codecs are noted+deferred. Read-only on game files; decoded files + DB rows are the only writes. Pass either db_path (the .autopsy.db) or game_path (whose default .autopsy/ DB is used). Optional out_dir overrides where decoded files land; asset_kinds restricts to textures, meshes, or audio.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional cap on textures decoded (for quick smoke runs).
db_pathNoPath to the <game>.autopsy.db produced by unpack_game.
out_dirNoOptional output dir for decoded PNGs. Defaults to <db_dir>/decoded/.
game_pathNoGame install dir; its default .autopsy/<name>.autopsy.db is used if db_path is omitted.
asset_kindsNoWhich asset kinds to decode. 'all' (default) does textures + meshes + audio; restrict with 'textures', 'meshes', or 'audio'.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it declares 'Read-only on game files; decoded files + DB rows are the only writes', spells out exactly which formats are produced vs noted-and-deferred (compressed meshes, FSB5-wrapped Vorbis), and describes the resS-sidecar streaming. This is rich disclosure beyond what a structural field would convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then conversion details, then write semantics and parameters. It is information-dense and every clause earns its place, though the arrow-chain syntax is slightly heavy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations exist, yet the description covers the input modes, conversion outputs, side effects (decoded files + DB rows), and deferral behavior, and notes the decoded=1/decoded_path DB update. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: the db_path/game_path either-or relationship and the override behavior of out_dir and asset_kinds. It clarifies the mutual-exclusion semantics that the schema only implies per-parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Asset Decoder') and details exactly what it does: reads a PR-1 autopsy catalog, opens Unity .assets containers, and converts members to standard formats. The per-type conversion list (Texture2D -> PNG, Mesh -> glTF, AudioClip -> WAV/OGG) makes its identity unmistakable against siblings like read_asset and list_assets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains how to invoke it ('Pass either db_path or game_path') and the effect of asset_kinds, but gives no explicit when-to-use/when-not guidance and names no alternative tool. The usage context is implied by 'Universal Asset Decoder' rather than stated, so it lands at implied-usage rather than a routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decompile_methodB

Decompile a single method to C# source code.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesFull type name
method_nameYesMethod name
parameter_countNoParameter count to disambiguate overloads

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but only mentions the output format (C# source). It does not disclose whether the assembly must already be loaded, whether the operation is read-only, or any other behavioral traits like permissions or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that efficiently states the tool's purpose with zero waste. It is appropriately sized for a simple decompilation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the return value (C# source code) and the tool's core action. However, it omits contextual prerequisites, most notably that an assembly must be loaded before use, which is important for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all four parameters. The description adds no additional meaning beyond what the schema provides, making the baseline score of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Decompile') and resource ('a single method') with the output format ('C# source code'), clearly distinguishing it from sibling tools like decompile_type (whole type) and get_method_il (IL). An agent can tell exactly what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., decompile_type, get_method_il, inspect_type) or any prerequisites such as loading an assembly first. The description only states the purpose without usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decompile_typeC

Decompile a type to full C# source code.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesFull type name (Namespace.ClassName)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the output form (full C# source code), which is useful, but says nothing about preconditions (e.g., whether the assembly must be loaded first), failure modes for obfuscated or unresolvable types, or size/limits of the emitted source.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste, but it is arguably too terse for a two-parameter tool where a clause on prerequisites or output scope would be inexpensive to add.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema fully covers the two inputs and the description states the return form, so the core is present. However, with no annotations and no output schema, the relationship to load_assembly and error behavior for unresolvable types remain unstated, leaving an agent to guess at prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (assembly key, full Namespace.ClassName). The description adds no format or syntax detail beyond what the schema supplies; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (decompile) and resource (a type), plus the output form (full C# source code). It implicitly distinguishes itself from the sibling decompile_method, which targets a single method rather than a type, though it does not name that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no preconditions, and no routing to alternatives. An agent must infer from the name that this is the whole-type counterpart to decompile_method and inspect_type.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_engineB

Scan a game directory and detect what engine it uses (Unity, Unreal, Godot, Java, etc.). Returns engine type, runtime, and primary assembly path.

ParametersJSON Schema
NameRequiredDescriptionDefault
game_pathYesFull path to the game directory

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It does disclose the return values (engine type, runtime, primary assembly path), which is useful, but says nothing about permissions, read-only nature explicitly, or failure modes for unrecognized directories.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the primary action and then the outputs. No wasted filler, though the return-value sentence partly compensates for the absent output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter detection tool with no output schema, the description covers the action and expected outputs adequately. The main gap is usage guidance relative to sibling detection tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single parameter, so the schema already documents that game_path is the full path to the game directory. The description adds no additional syntax or format detail beyond that baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Scan... and detect') and resource ('game directory'), and enumerates engine types (Unity, Unreal, Godot, Java). This makes the tool's job clear, though it doesn't explicitly contrast with siblings like analyze_file_format or read_unity_assets_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no alternatives to consider. It only describes what the tool does, leaving the agent to infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_networkingC

Detect networking frameworks and flag unsafe-to-patch methods.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it never states that this is a read-only detection pass, whether it requires a loaded assembly, or what the flags look like. It hints at output ('flag unsafe-to-patch methods') but gives no real behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no padding and the core action front-loaded. It is efficient, though borderline terse for a tool with an unexplained dual purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description omits what a networking framework detection covers, what constitutes an unsafe-to-patch method, and whether the assembly must be pre-loaded. An agent would need to guess at prerequisites and result shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter ('assembly') with 100% schema description coverage, so the schema already documents it and the description adds nothing. Baseline 3 applies when the schema does the parameter work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a specific verb (detect) and resource (networking frameworks) and adds a second capability (flagging unsafe-to-patch methods). An agent can grasp what it does, though it does not distinguish itself from the analogous detect_engine sibling or clarify how the two tasks relate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to choose instead (e.g., detect_engine, list_patchable_methods). The only guidance is implied by the phrase 'unsafe-to-patch', leaving the agent to infer its role in a patch workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff_assembliesC

Compare two assembly versions to show changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 200)
show_methodsNoShow method-level changes (default: true)
new_assembly_pathYesPath to new version DLL
old_assembly_pathYesPath to old version DLL

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and falls short: it does not state that the operation is read-only, whether it is expensive on large assemblies, or how results are truncated (the limit param defaults to 200). 'Show changes' is the only behavioral hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. It is efficient, though arguably too terse to be maximally useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a comparison tool with no output schema, the description should indicate the shape of the result (e.g. added/removed/changed types and methods) and note the limit behavior. Neither is present, leaving the caller unsure what they will receive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (paths, limit, show_methods) are already documented in the schema. The description adds no extra meaning such as path format expectations or what show_methods=false suppresses, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb+resource (compare two assembly versions) and states the outcome (show changes), so an agent knows what it does. However, it does not distinguish itself from nearby siblings such as compare_binaries_detailed, compare_signatures, or jar_diff, which a caller must disambiguate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over the other diff/compare tools in the sibling list. No prerequisites (e.g. assemblies must be loaded or on disk) or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disassemble_functionA

Disassemble a function starting at an RVA in a native PE (x86-64). Linear sweep from the entry RVA to the first terminal instruction (ret/iret) or max_bytes, whichever comes first — a PR-3.1 boundary heuristic (CFG-accurate bounds arrive with the call-graph in PR-3.2), so the function EXTENT is tagged provenance=inferred while the decoded bytes themselves are verified. Read-only on the binary. Writes a 'function' row into The Model when db_path/game_path is supplied and the binary is catalogued.

ParametersJSON Schema
NameRequiredDescriptionDefault
rvaYesEntry RVA of the function.
db_pathNoOptional .autopsy.db for writeback.
game_pathNoOptional game dir; default .autopsy DB used if db_path omitted.
max_bytesNoSafety cap on sweep length (default 4096).
binary_pathYesPath to the PE (.exe/.dll).

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses read-only access on the binary, the conditional Model writeback when db_path/game_path is supplied, and the provenance distinction (extent inferred vs. decoded bytes verified). It omits error/failure behavior and the concrete return shape, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded, but the sentence is dense and includes internal roadmap references (PR-3.1, PR-3.2) that add noise for an agent without aiding invocation. Every clause about provenance is useful; the PR tags are not.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-trivial disassembler with no output schema and no annotations, the description covers the boundary heuristic, provenance tagging, read-only nature, and writeback trigger. It stops short of describing the returned data structure, which the absence of an output schema would ideally have motivated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline is 3, but the description adds real meaning: it clarifies that max_bytes bounds the sweep and, crucially, explains the conditional write semantics of db_path/game_path ('when db_path/game_path is supplied and the binary is catalogued') beyond the schema's terse 'for writeback'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (disassemble) and resource (a function at an RVA) with scope constraints (native PE, x86-64). The boundary-heuristic detail implicitly separates it from a raw range disassembler like disassemble_range, but no sibling is named explicitly, so differentiation is inferential rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the internal algorithm but never says when to choose this over disassemble_range or analyze_pe_full, nor any preconditions (e.g., binary must be a valid PE, catalogued for writeback). Usage is only implied by the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disassemble_rangeA

Disassemble an arbitrary byte range of a native PE (x86-64) starting at an RVA. Resolves the RVA to its section and file offset, reads ONLY that slice (read-only on the binary), and decodes it with Capstone. Returns the instruction listing. If db_path or game_path is given AND the binary is catalogued in that autopsy DB, the decoded instructions are also summarized into The Model as a 'function' row spanning the range (provenance=verified for the bytes). Pure read-only when no DB is supplied.

ParametersJSON Schema
NameRequiredDescriptionDefault
rvaYesRelative virtual address to start at (e.g. 0x1000).
lengthNoNumber of bytes to disassemble (default 256, clamped to the section).
db_pathNoOptional .autopsy.db for writeback.
game_pathNoOptional game dir; its default .autopsy DB is used if db_path is omitted.
binary_pathYesPath to the PE (.exe/.dll) to disassemble.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses read-only behavior on the binary, the conditional Model writeback, and clarifies 'Pure read-only when no DB is supplied.' It does not cover error/edge behavior (out-of-range RVA, clamping details beyond the schema), so it stops short of fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the core action and the read/write consequence, then the conditional side effect. Mostly efficient, though the final 'Pure read-only' sentence slightly restates the earlier 'read-only on the binary'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description covers return value ('instruction listing') and the side-effect contract adequately. It leaves error handling and behavior when the RVA/section is invalid unspecified, but the essential call-time information is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all five parameters with examples and defaults. The description adds relational context for db_path/game_path (writeback triggers only when the binary is catalogued), but this is largely redundant with the schema's own descriptions of those params. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: disassemble an arbitrary byte range of a native PE (x86-64) starting at an RVA. The 'arbitrary byte range ... starting at an RVA' framing implicitly distinguishes it from the sibling disassemble_function (which targets a named function), so an agent can tell them apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the condition under which writeback occurs ('If db_path or game_path is given AND the binary is catalogued'), which is useful context, but it never explicitly says when to choose this over disassemble_function or rva_to_offset. Usage is left implied by the RVA-based framing rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_dll_classesA

Extract class names from a .NET DLL using stream-based analysis. Works with very large files.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPath to DLL file
max_classesNoMax classes to return (default 200)
search_termsNoFilter by these terms (case-insensitive)

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one real behavioral trait: stream-based analysis that tolerates very large files. It says nothing about error handling on non-.NET/invalid binaries, output ordering, or whether the scan is exhaustive or truncated, leaving meaningful behavioral gaps for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and resource, with the second sentence adding a distinct constraint. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description is the only source for behavior, yet it never states the return shape (e.g., flat list of fully-qualified names vs. nested), ordering, or failure mode on non-.NET files. It is adequate for a simple three-parameter read tool but leaves those gaps unfilled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so file_path, max_classes (default 200), and search_terms (case-insensitive filter) are already fully documented in the schema. The description adds no parameter-level detail such as filter matching semantics against namespaces or what happens when max_classes truncates results, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ("Extract") and resource ("class names from a .NET DLL"), so an agent knows exactly what it produces. It doesn't differentiate itself from close siblings like list_types or analyze_dll_structure, which likely also enumerate types from assemblies, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase "Works with very large files" implicitly signals when this tool is preferable (huge DLLs where stream-based parsing matters), but there is no explicit when-to-use, when-not-to-use, or named alternative such as list_types or analyze_dll_structure. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_stringsC

Simple string extraction from a binary file.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional filter
file_pathYesPath to binary file
min_lengthNoMinimum string length (default 4)

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only extraction but does not state whether it modifies the file, how output is returned, whether results are deduplicated, or how encoding/ordering is handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, so it is concise, but it is arguably under-specified rather than efficient. 'Simple' is a vague qualifier that consumes space without adding information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a sibling variant, three parameters, and no output schema, the description omits return format, filter semantics, and the simple-vs-advanced distinction. An agent cannot call it correctly relative to its alternatives from this text alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (filter, file_path, min_length) are already documented in the schema. The description adds no meaning beyond this, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource: string extraction from a binary file, which is more specific than a tautology. However, it does not distinguish itself from the sibling 'extract_strings_advanced', leaving the agent unsure which variant to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'Simple' hints that a more capable variant exists, but no explicit when-to-use guidance, prerequisites, or alternative (extract_strings_advanced) is named. The agent must infer the routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_strings_advancedC

Extract ASCII and UTF-16 strings from a binary file with filtering.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional filter (case-insensitive substring match)
encodingNoEncoding: ascii, utf16, both (default: both)
file_pathYesPath to the binary file
min_lengthNoMinimum string length (default 4)
max_resultsNoMax results (default 500)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says nothing about return format, ordering, truncation behavior when max_results is hit, or performance on large binaries, all of which matter for a scanning tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the core action front-loaded and no wasted words. It is arguably under-specified rather than bloated, which is a mild cost but not a structure problem.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and five parameters, the description should explain returns and the distinction from 'extract_strings'. It does neither, leaving the definition thin for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all five parameters (filter, encoding, file_path, min_length, max_results) are documented in the schema itself. The description's mention of ASCII/UTF-16 and filtering mirrors the encoding and filter params but adds no syntax or format detail beyond the schema, so the baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Extract ASCII and UTF-16 strings from a binary file') with the added filtering scope. It is clear, but it never differentiates itself from the sibling 'extract_strings' despite the 'advanced' name, leaving the agent unable to tell which of the two to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative routing is given. With a near-identical sibling ('extract_strings') present, the absence of any guidance on when 'advanced' is preferred is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_blueprints_of_typeB

Search for all Blueprint assets inheriting from a parent class.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 50)
path_filterNoOptional path filter
parent_classYesParent class name

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Search' implies a read-only operation, but it discloses nothing about return format, ordering, pagination beyond the limit parameter, or whether results include nested subclasses transitively.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler and the scoping constraint positioned for quick scanning. It is efficient, though its brevity edges toward under-specification rather than being a model of conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should explain what the search returns and how parent_class matching works, but it does neither. For a search tool whose only output documentation would have to live here, this leaves meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (limit, path_filter, parent_class) are already documented in the schema. The description adds no format or constraint detail beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Search) and resource (Blueprint assets) with a clear scope qualifier (inheriting from a parent class). An agent can distinguish it from read_blueprint or list_assets by the inheritance filter, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The inheritance scope implicitly tells the agent when this tool applies (finding blueprint subclasses of a given class). However, it gives no explicit guidance versus close neighbors like get_inheritance_tree, list_types, or search_assets, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_gameplay_valuesC

Search for gameplay-related values (health, damage, speed, money, etc.) across the game code. Works with any engine by using the appropriate search tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
search_termsYesTerms to search for (e.g. ["Health", "Damage", "Speed", "Money"])
assembly_pathYesPath to the game assembly/binary to search

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It does not state that this is a read-only operation, nor describe result format, matching semantics, or performance across large assemblies. The engine-agnostic note is a minor behavioral hint but far from sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core purpose front-loaded and no filler. Slightly thin rather than padded, so it uses its space efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema and no annotations, the description should explain what a match looks like (offset, type, value) and how results are scoped. It covers the what and the cross-engine angle but omits the return shape, so it is minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both search_terms and assembly_path with examples. The description's parenthetical examples (health, damage, speed, money) merely echo the schema rather than adding format or matching semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (gameplay-related values such as health, damage, speed, money) across game code, which is more semantic than raw scanning. However, it does not differentiate itself from the many sibling search tools (search_code, pattern_scan, search_binary_pattern), leaving the agent to infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is that it 'works with any engine by using the appropriate search tools,' which hints at cross-engine applicability but never states when to use this over the obvious alternatives like search_code or pattern_scan. No exclusions or selection conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_referencesC

Find all references to a method or field.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 100)
assemblyYesAssembly key
type_nameYesType containing the member
member_nameYesMember name to find references to

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden but only implies a read. It does not disclose search scope (single assembly vs. all loaded assemblies), whether results are capped by the limit parameter, or what the response looks like. For a search tool with zero annotation coverage this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler or redundancy. It is efficient, though the extreme brevity contributes to the missing context noted in other dimensions rather than being a flaw in structure itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description would need to explain return values, but it does not. It also omits search scope and limit behavior. For a 4-parameter reference-search tool with no annotations and no output schema, the definition is too thin to fully prepare an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (assembly, type_name, member_name, limit) are already documented in the schema, which sets the baseline at 3. The description adds only marginal value by clarifying that the target member may be a method or field, which the schema does not state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Find) and resource (references to a method or field), so an agent knows exactly what operation is performed. However, it offers no differentiation from related siblings such as search_code, pattern_scan, or inspect_type, which an agent could easily confuse for reference-finding.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus search_code, pattern_scan, or inspect_type, nor any stated prerequisites (e.g., that the assembly must be loaded). The description is purely a restatement of the operation with no routing or context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_renamed_typesB

Find types likely renamed after a game update.

ParametersJSON Schema
NameRequiredDescriptionDefault
min_similarityNoMin similarity threshold 0-1 (default: 0.6)
new_assembly_pathYesPath to new version DLL
old_assembly_pathYesPath to old version DLL

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing. The word 'likely' hints at fuzzy/heuristic matching, but there is no statement about read-only safety, what the matching is based on, runtime cost on large assemblies, or what the result set looks like. A single sentence is inadequate for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler, and the key concept (renamed types) appears immediately. It is efficient, though for a heuristic tool of this complexity the brevity borders on under-specification rather than tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, no annotations, and the description never explains what a result entry contains or how candidates are ranked beyond the similarity threshold. For a heuristic analysis tool whose output shape is entirely undocumented, the agent is left guessing what it will receive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents min_similarity with its range and default plus both assembly paths. The description adds no parameter meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (find) and resource (types) with a qualifier that scopes the intent: 'likely renamed after a game update'. That is enough for an agent to know this is a heuristic rename-detection tool. It does not, however, distinguish itself from siblings like diff_assemblies or compare_signatures, which an agent might reasonably consider for the same task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'after a game update' implies the scenario in which to reach for this tool, which is more than nothing. But it offers no explicit when-not conditions and never names an alternative such as diff_assemblies or compare_signatures, so the agent must infer the boundary itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_steam_gamesB

Search configured Steam library paths for installed games. Optionally filter by name.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional name filter (case-insensitive partial match)

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It reveals that search targets configured Steam library paths, but does not state whether the operation is read-only, whether it can be slow or require Steam to be configured, or what side effects (if any) exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste, front-loading the core purpose before the optional parameter. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, single-parameter tool with no output schema and no annotations, the description covers the essential purpose and parameter. It could still hint at return shape or read-only nature, but the core information needed to invoke it is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the sole parameter already documents itself as an optional case-insensitive partial match. The description only restates 'Optionally filter by name,' adding no extra parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search'), resource ('installed games'), and scope ('configured Steam library paths'). It is clearly distinguishable from all sibling tools, none of which deal with listing or finding Steam games.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what the tool does but gives no guidance on when to use it versus alternatives. It only mentions an optional parameter, not when a search is appropriate or what prerequisites exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

game_statusA

Get the current session state including loaded game, engine type, cached analysis data, and mod build status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the returned data categories, which is useful behavioral context for a status tool, but it does not explicitly state that it is read-only, has no side effects, or requires no authentication. For a no-parameter getter the risk is low, but the gap prevents a higher score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that lists the included state components with no filler. Appropriate size for a zero-parameter status tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return content, and it does list the main state categories. It is complete enough for a simple status getter, though it could clarify the output format or the distinction from unreal_game_status.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document; baseline 4 applies. The description does not need to add parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Get') and resource ('current session state'), and it enumerates the included components (loaded game, engine type, cached analysis data, mod build status). However, it does not distinguish this tool from the sibling unreal_game_status or detect_engine, so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives, and no prerequisites. The agent receives no help deciding between this tool and siblings like unreal_game_status or detect_engine.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_harmony_patchC

Generate correct Harmony patch class for a specific method.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesTarget type name
patch_typeNoPatch type: prefix, postfix, both (default: both)
method_nameYesTarget method name
parameter_countNoParameter count for overload disambiguation
patch_class_nameNoCustom name for the patch class

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden: it does not say whether a file is written to disk or source text is returned, whether the target assembly must already be loaded, whether an existing patch class is overwritten, or how errors surface. 'Correct' is asserted but not explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It is efficient, though the brevity contributes to the missing guidance rather than being a virtue of precision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter code-generation tool with no annotations and no output schema, the one-line description is insufficient: the agent cannot tell what is produced (file vs. code), where it lands, or what prerequisites exist before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (including patch_type default 'both' and parameter_count for overload disambiguation) are documented in the schema itself. The description adds no parameter meaning beyond what the schema provides, which is the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (generate) and resource (Harmony patch class) scoped to a specific method, so the agent knows the outcome. It does not, however, distinguish this from siblings like validate_patch_target, list_patchable_methods, or generate_plugin, which also operate on patch targets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no preconditions (e.g. whether the assembly must be loaded first via load_assembly), and no mention of alternatives such as validate_patch_target or compile_plugin. The agent must infer the entire workflow position of this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_pluginC

Generate complete BepInEx 5 plugin template with Harmony patching.

ParametersJSON Schema
NameRequiredDescriptionDefault
authorNoAuthor name
versionNoVersion string (default 1.0.0)
descriptionNoPlugin description
plugin_guidYesPlugin GUID (e.g. com.author.pluginname)
plugin_nameYesPlugin class name
include_configNoInclude config system
include_harmonyNoInclude Harmony setup
managed_directoryNoPath to game Managed directory

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full behavioral burden, yet it discloses only what is generated. It says nothing about where files are written to disk, whether it overwrites existing files, whether the codebase must be loaded first, or what the agent should do afterwards (e.g. compile_plugin).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero redundancy. It states the action and the artifact immediately with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter generation tool with no annotations and no output schema, the description is too thin. It does not convey side effects, filesystem behavior, output location, or how it relates to the surrounding mod-build pipeline, leaving meaningful gaps for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters, and the baseline is 3. The description adds no additional parameter meaning, such as which optional flags alter the generated output structure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Generate) and resource (BepInEx 5 plugin template with Harmony patching), so the agent knows exactly what artifact it produces. It does not, however, distinguish itself from close siblings like scaffold_mod or gorebox_generate_mod, which fill a similar template-generation role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to reach for this tool versus generate_harmony_patch, scaffold_mod, compile_plugin, or the jar/gorebox generation tools. There are no prerequisites, conditions, or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_class_hierarchyC

Get inheritance chain for a Blueprint class.

ParametersJSON Schema
NameRequiredDescriptionDefault
asset_pathYesBlueprint asset path
max_levelsNoMax inheritance levels (default 10)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Get' implies a read, but nothing states whether the walk is bounded, whether it errors on non-Blueprint assets, or how the chain is represented; only the schema's max_levels hint suggests truncation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though its brevity reflects under-specification rather than disciplined trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and a trivially small parameter set, the description should at least convey what the returned chain contains and how depth limiting affects it. What is present is technically accurate but too thin to guide correct invocation confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both asset_path and max_levels (with its default of 10) are already documented in the schema. The description adds no format or naming detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Get inheritance chain for a Blueprint class' — so the operation is unambiguous. It does not, however, differentiate itself from the sibling get_inheritance_tree, leaving the agent to guess which of the two inheritance tools to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the near-identical sibling get_inheritance_tree. The agent receives no criteria for choosing this tool over that one or over inspect_type/decompile_type.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_inheritance_treeB

Get full inheritance tree for a type (ancestors and descendants).

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesFull type name

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. "Get" implies a safe read, but nothing is said about depth limits, whether the tree is transitive or direct-only, cycle handling for recursive types, or whether the assembly must be loaded beforehand.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and resource, then disambiguates the scope in parentheses. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no output schema, the description is minimally adequate: it names the resource and the two directions traversed. It omits prerequisite state (loaded assembly), result shape, and relationship to the near-duplicate get_class_hierarchy sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both required parameters (assembly, type_name) are already documented in the schema. The description's "full type name" phrasing loosely echoes type_name but adds no new syntax, format, or constraint information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Get full inheritance tree for a type") and clarifies the scope with "(ancestors and descendants)", which is more precise than the sibling get_class_hierarchy. It does not explicitly distinguish itself from that sibling or from inspect_type, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over get_class_hierarchy, inspect_type, or list_types, nor any stated prerequisites (e.g., assembly must be loaded first). The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_method_ilC

Get raw IL bytecode for a method.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesFull type name
method_nameYesMethod name

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does little: it implies a read operation but never states safety, whether the assembly must first be loaded via load_assembly, how large the output might be, or how the raw IL is formatted. 'Raw IL bytecode' is the only hint about the return value, which is thin for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is efficient, though its brevity edges toward under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter lookup tool with no output schema and no annotations, the description is minimally adequate but leaves key gaps: how the IL is returned, whether the assembly must be loaded first, and how it differs from the decompile/disassemble siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (assembly, type_name, method_name) are already documented in the schema and the description adds no further semantics such as key format or name qualification rules. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (raw IL bytecode for a method), so the agent knows it returns IL rather than decompiled source. It does not, however, distinguish itself from siblings like decompile_method, disassemble_function, or jar_search_bytecode, which an agent must infer on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all: nothing tells the agent to prefer this over decompile_method when it needs raw IL versus C#, nor what conditions (e.g., loaded assembly, patch workflow) make this the right call. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_serialized_fieldsC

Get all serialized fields for a MonoBehaviour or ScriptableObject.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesFull type name

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It implies a read but never states that it is non-destructive, what happens if the type is not a MonoBehaviour/ScriptableObject, whether the assembly must already be loaded, or what form the fields come back in. For a query tool with zero annotation coverage this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core action and target-types front-loaded, and no filler. It is efficient, though there is very little content to be concise about.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, two-required-parameter read tool with a fully documented schema and no output schema, the description is minimally sufficient. It still omits prerequisites and the nature of the returned field data, which an agent would need to use and chain the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both required parameters ('Assembly key', 'Full type name') are documented in the schema, so the baseline of 3 applies. The description adds only the target-type constraint (MonoBehaviour/ScriptableObject), not format or lookup rules for assembly/type_name beyond what the schema says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Get all serialized fields') and restricts the target to MonoBehaviour or ScriptableObject, which separates it from generic siblings like inspect_type or decompile_type. It lacks explicit naming of which sibling to prefer when the agent wants broader type metadata, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., whether the assembly must be loaded first), and no mention of alternatives such as list_monobehaviours, inspect_type, or decompile_type. The agent must infer that this is the narrow serialization-inspection path rather than the general type-inspection path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gorebox_generate_discovery_modB

Generate a GoreBox Lua mod that discovers all available API functions, globals, tables, and callbacks. Writes results to a dump file in the mod folder. Use this when you need to learn what Lua API is available in a game that uses Lua scripting.

ParametersJSON Schema
NameRequiredDescriptionDefault
mod_nameNoName for the discovery mod (default: APIDisco)
mods_dirYesPath to the game Mods directory (e.g. C:/Users/.../GoreBox/Mods)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It does disclose a side effect — results are written to a dump file in the mod folder — which is genuinely useful. But it omits auth/permission needs, whether it overwrites an existing mod, and how the generated mod is deployed or run, leaving meaningful behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action front-loaded and the conditional trigger placed last. No filler or repetition; efficiency is good, though it could be marginally tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with real side effects and no annotations or output schema, the description should cover the end-to-end flow — where the dump lands and that gorebox_read_api_dump consumes it. It explains what is generated but not how the output is used, leaving the agent short of what it needs to act on the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (mods_dir, mod_name) are already documented with types and defaults. The description adds no syntax, format, or path-caveat detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: generating a Lua mod that enumerates API functions, globals, tables, and callbacks. This is precise. However, it does not distinguish itself from the sibling gorebox_generate_mod, which an agent could easily confuse with this specialized discovery generator.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this when you need to learn what Lua API is available in a game that uses Lua scripting" gives a use condition, but it largely restates the purpose. It never names the natural follow-up tool (gorebox_read_api_dump) or any alternative, leaving the workflow to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gorebox_generate_modC

Generate a GoreBox-compatible Lua mod with proper info.json and main.lua. Creates a ready-to-use mod in the Mods directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoVersion string (default 1.0.0)
lua_codeYesThe Lua script code for main.lua
mod_nameYesName for the mod folder
mods_dirYesPath to the game Mods directory
is_pluginNoWhether this is a plugin mod (default true)
safe_modeNoWhether to run in safe mode (default false)
descriptionNoMod description
display_nameNoDisplay name in the mod browser

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses only that a mod is written into the Mods directory; it says nothing about overwrite behavior, whether existing mod folders are clobbered, required permissions, or side effects on the game installation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the outcome and then the artifact location. No filler, though the second sentence is somewhat redundant with the first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an eight-parameter file-generating tool with no annotations and no output schema, the description omits too much: no success/failure signal, no overwrite semantics, no distinction from sibling generators, and no indication of what the caller should do with the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all eight parameters (including is_plugin and safe_mode defaults) are already documented in the schema. The description adds no syntax, validation, or default detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Generate a GoreBox-compatible Lua mod") and names the two files produced (info.json, main.lua). It is clear what the tool produces, though it never names or differentiates itself from the close sibling gorebox_generate_discovery_mod.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives. An agent cannot tell from the description how this differs from gorebox_generate_discovery_mod or the generic scaffold_mod, both of which appear to produce mod scaffolding.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gorebox_list_modsB

List all installed GoreBox mods with their info.json contents and file structure.

ParametersJSON Schema
NameRequiredDescriptionDefault
mods_dirYesPath to the game Mods directory

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose useful output behavior (returns info.json contents plus file structure), but says nothing about read-only safety, required permissions, or how incomplete/missing mod directories are handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with zero filler; the verb, resource, and returned content are all packed into a single readable clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter list tool with no output schema, the description covers the basic return surface (info.json and file structure) but is thin on the shape of results and error behavior. It is adequate but leaves gaps an agent might want when interpreting output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage ('Path to the game Mods directory'), so the schema already documents it. The description adds no format, path-style, or defaulting details beyond the schema, which is the expected baseline when schema coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (List) and resource (installed GoreBox mods) and clarifies the scope by stating it returns info.json contents and file structure. It is inherently distinguishable from generation siblings like gorebox_generate_mod, but it never explicitly contrasts with them or with any related listing tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternative tools. The sentence only states what the tool does, leaving the agent to infer when listing mods is appropriate versus other gorebox_* or asset-inspection tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gorebox_read_api_dumpA

Read and parse an API discovery dump file generated by gorebox_generate_discovery_mod. Returns structured information about all discovered globals, functions, tables, and callbacks.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional filter string to match entries (case-insensitive)
dump_pathYesPath to the dump.txt file

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Read and parse' correctly signals a non-mutating operation, and the returned data categories are disclosed. It does not state behavior on a missing/invalid dump_path, whether results are cached, or size/rate considerations for large dumps, leaving real gaps for a zero-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with the core action front-loaded and the return contents following immediately. No filler or redundancy, though it is a fairly minimal statement of behavior rather than a maximally information-dense definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no output schema and full schema coverage, the description supplies what an agent needs: the action, the input source, and the shape of the returned data. Remaining gaps (error behavior, filter semantics) are minor at this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both dump_path and filter fully described in the schema, so the baseline is 3. The description adds nothing beyond the schema about parameter formats or the semantics of the case-insensitive filter, so no credit above baseline is earned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair ('read and parse') and resource ('API discovery dump file'), and names the upstream generator (gorebox_generate_discovery_mod) so the agent can see this is the consumer of that tool's output. It also enumerates what is returned (globals, functions, tables, callbacks), making the scope unambiguous. It does not explicitly contrast against other analysis siblings like inspect_type or pattern_scan, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The reference to the generating tool implies the intended workflow (generate a discovery dump, then read it), which is useful sequencing context. However, there is no explicit when-to-use vs. when-not, no guidance on the optional filter parameter's purpose, and no stated prerequisite that the dump file must already exist. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hex_readC

Read raw bytes from a file at a specific offset.

ParametersJSON Schema
NameRequiredDescriptionDefault
lengthNoNumber of bytes to read (default 256)
offsetYesByte offset to start reading
file_pathYesPath to the file

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read-only operation but never states safety, whether offsets beyond EOF error or truncate, how much is returned when fewer than 'length' bytes remain, or the return encoding (hex string vs raw bytes). For a binary-reading tool with zero annotation coverage, this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the operation and its two key inputs are stated immediately. It is terse to the point of under-specification, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter read tool with no output schema, the essentials of the operation are conveyed. However, the return representation (hex dump, raw bytes, base64) and out-of-range behavior are unstated, and with no output schema to fall back on the agent must guess what comes back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents file_path, offset, and the length default of 256. The description only restates 'offset' and adds nothing about units, bounds, or the default length, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read'), resource ('raw bytes from a file'), and scope ('at a specific offset'), so the agent knows it is a byte-level read primitive. It does not explicitly differentiate itself from siblings like hex_write, hex_replace, or jar_hex_view, but the read/offset framing is distinctive enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and names no alternatives. With hex_search, hex_write, hex_replace, and jar_hex_view all present, the agent gets no help choosing between them; it must infer from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hex_replaceC

Replace bytes at a specific offset or replace a pattern throughout the file.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNoSpecific offset to replace at (if provided, ignores search_hex)
file_pathYesPath to the file
search_hexNoHex pattern to find
replace_hexYesHex pattern to replace with (must be same length)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden for a mutation tool. It never states that the file is modified in place, whether a backup is made, whether changes are reversible, what permissions are needed, or what happens on a pattern mismatch or length mismatch — significant gaps for a destructive write.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the two modes front-loaded and no filler. It is appropriately sized, though the absence of any operational caveat means it is lean to the point of under-informing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with no annotations and no output schema, the description is too thin. It omits in-place editing semantics, error behavior, and the consequences of pattern vs. offset mode, leaving an agent without enough context to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the offset-overrides-search_hex precedence and the equal-length constraint. The description adds only the idea that search mode acts 'throughout the file' (all occurrences), a small increment over the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (replace) and resource (bytes/hex) and distinguishes its two operating modes: offset-targeted vs. pattern-wide. It does not, however, differentiate itself from closely related siblings like hex_write, hex_search, or jar_hex_edit, so 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives. The agent must infer from the description alone that offset mode and search mode are mutually exclusive; nothing routes it between this tool and hex_write or pattern_scan.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hex_writeC

Write raw bytes to a file at a specific offset.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetYesByte offset to write at
hex_dataYesHex string to write (e.g. "90 90 90" or "909090")
file_pathYesPath to the file

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does not say whether existing bytes are overwritten, whether the file grows past EOF, what happens on a failed write, or what permissions/locking are required. For a mutation tool with zero annotation coverage this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though the extreme brevity is part of why behavioral and usage gaps remain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, destructive-capable tool with no annotations and no output schema, the description omits overwrite semantics, boundary behavior, and failure modes. The schema covers parameters well, but the description is not complete enough for an agent to invoke this safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented with format examples. The description's phrase 'at a specific offset' and 'raw bytes' merely restates the schema, adding no syntax or constraint detail beyond it. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (write), resource (raw bytes to a file), and locus (specific offset), so the core action is unambiguous. However, it does not differentiate itself from close siblings like hex_replace, jar_hex_edit, or jar_hex_view, which an agent must distinguish by name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no named alternative. The agent gets no help deciding between this and hex_replace, which appear to overlap on the same file/byte surface.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_typeA

Inspect type structure (fields, properties, methods) without decompiling method bodies.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesFull type name

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It signals a read-only inspection and adds a useful scope boundary (method bodies are not decompiled), but says nothing about permissions, cost, or what the output looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence that leads with the action and scope boundary. Every clause earns its place and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool with no output schema and no annotations, the description covers purpose and scope but does not sketch the return shape (how fields/methods are presented) or note any constraints on the assembly/type identifiers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (assembly key, full type name), so the schema already documents both parameters. The description adds no syntax or format detail beyond that, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (inspect) and resource (type structure) and enumerates what is returned (fields, properties, methods). The clause 'without decompiling method bodies' implicitly separates it from the sibling decompile_type, though it never names that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: you would reach for this when you need type shape but not method bodies. There is no explicit when-to-use statement, no mention of decompile_type as the alternative when bodies are needed, and no prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_add_fileC

Add a file from disk into the JAR.

ParametersJSON Schema
NameRequiredDescriptionDefault
jar_pathYesDestination path within JAR
overwriteNoOverwrite if exists (default: true)
session_idYesSession ID
source_pathYesSource file on disk

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does not meet it. It doesn't state whether a session/JAR must already be open, that existing entries are overwritten by default (schema's overwrite=true), or whether the operation is reversible. For a mutation tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It is efficient, though its brevity borders on under-specification rather than true economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and a required session_id, the description is too thin. It omits session lifecycle expectations, overwrite behavior, and success/failure semantics, leaving real gaps an agent needs filled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (session_id, jar_path, source_path, overwrite) are already documented in the schema. The description adds no syntax or format detail beyond it, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Add") and resource ("a file from disk into the JAR"), and the phrase "from disk" implicitly distinguishes it from the sibling jar_add_file_content. It stops short of naming that alternative, so the differentiation requires the agent to infer it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the obvious alternative jar_add_file_content (which presumably adds content rather than a disk file). The agent must guess which of the two file-adding tools to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_add_file_contentC

Write text content as a new file in the JAR.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesFile content
encodingNoEncoding (default: utf-8)
jar_pathYesDestination path within JAR
session_idYesSession ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not say what happens if the file already exists (overwrite vs error), whether the session must be open, or whether the change is reversible or persisted to disk on jar_repack.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though its brevity edges toward under-specification rather than tight conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no annotations and no output schema, the description omits the behavioral details an agent needs: overwrite semantics, session lifecycle, and encoding interaction. The name and schema cover the what, but the how and the risks are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (content, encoding, jar_path, session_id) are already documented in the schema. The description adds no syntax, format, or path conventions beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Write) and resource (text content as a new file in the JAR), which lets an agent distinguish it from a read tool like jar_read_file. It does not, however, differentiate itself from the closely named sibling jar_add_file, leaving the choice between them ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no mention of alternatives. The existence of jar_add_file and jar_edit_file as siblings makes the absence of routing guidance a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_closeA

Close a JAR session and clean up temp files.

ParametersJSON Schema
NameRequiredDescriptionDefault
cleanupNoDelete temp files (default: true)
session_idYesSession ID

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the core effect (session termination plus temp-file cleanup), which is more than the name alone, but omits whether the session is invalidated, whether repeated closes fail, and that cleanup is destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the primary action and its side effect are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter, no-output-schema tool the essentials are present, but with zero annotation coverage the description should note session invalidation and the destructive nature of cleanup to fully equip an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters (session_id, cleanup) are already documented in the schema. The description adds no semantics about session_id validity or the cleanup default/override behavior beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (close) and resource (JAR session) plus the side effect of cleaning up temp files. The inverse relationship to jar_open makes it distinguishable without opening the schema, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied (close a session when finished). No statement of when to call it, prerequisites, or relationship to jar_open/jar_list_sessions, and no guidance on whether the cleanup path is optional in practice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_compile_javaC

Compile Java source code and optionally inject into JAR.

ParametersJSON Schema
NameRequiredDescriptionDefault
classpathNoAdditional classpath entries
class_nameYesFully-qualified class name
session_idYesSession ID
source_codeYesJava source code
java_releaseNoJava release target (default: 21)
auto_classpathNoAuto-add JAR to classpath (default: true)
inject_into_jarNoInject compiled class into JAR (default: true)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it omits key traits: the required session_id implies a pre-existing open JAR session (jar_open) that is never mentioned, and injecting a class into a JAR can overwrite existing entries — a destructive side effect left undisclosed. It also says nothing about permissions, failures, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words is well structured. It is arguably too terse for a seven-parameter mutation tool, but every stated clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a compile-plus-inject tool with seven parameters, no annotations, and no output schema, the description is not complete enough: it omits session prerequisites, the overwrite risk of injection, and the classpath/auto_classpath interaction, all of which an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters are documented in the schema, establishing the baseline of 3. The description adds only the notion that injection is optional, which loosely maps to inject_into_jar (default true) but provides no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Compile) and resource (Java source code) and notes the optional JAR injection, so the core action is unambiguous. It does not, however, distinguish itself from the related compile_plugin sibling or clarify how compilation output is delivered without injection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not guidance, and no alternatives are named despite siblings like compile_plugin, build_and_deploy, and jar_scaffold_mod. The phrase 'optionally inject into JAR' hints at a use case but leaves the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_detect_mod_infoC

Auto-detect mod loader, version, and structure from JAR contents.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only scan but doesn't state whether it mutates session state, persists results, requires an existing session, or how it handles unrecognized JARs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key action front-loaded and no filler. It is efficient, though almost too terse given the missing behavioral detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only detection utility this is nearly adequate, but with no output schema the description should at least hint at the returned shape (loader name, version fields). The prerequisite relationship to jar_open sessions is also unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (session_id) and schema coverage is 100%, so the schema already documents it fully. The description adds no syntax or format detail beyond that, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (detect) and a concrete resource (mod loader, version, structure) scoped to JAR contents. It is clearly distinguishable from generic siblings like detect_engine or jar_read_class, though it doesn't explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. It doesn't say whether a JAR must first be opened via jar_open/jar_list_sessions, which matters given the session_id parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_diffC

Show differences between original and modified file.

ParametersJSON Schema
NameRequiredDescriptionDefault
class_pathYesFile path within JAR
session_idYesSession ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it discloses almost nothing: no indication of whether the operation is read-only, what the output looks like, what 'modified' means, or what happens if there are no changes. One sentence is far too thin for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. It is efficient, though the brevity comes at the cost of substance rather than being fully earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should explain what the diff returns and how a session establishes the 'original' baseline. Neither is addressed, leaving the agent unable to predict the result of calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the schema already documents class_path and session_id. The description adds no additional meaning about how either parameter selects or scopes the diff, matching the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Show differences') plus resource ('original and modified file'), so the agent knows this is a diff operation. It does not say which 'original' or 'modified' refer to (session edits vs. on-disk) and does not distinguish itself from siblings like verify_patches or compare_binaries_detailed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the many alternative comparison tools in the sibling list. The agent must infer that a session must be open and edited before this is useful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_edit_class_constantsC

Edit numeric constants in .class file bytecode.

ParametersJSON Schema
NameRequiredDescriptionDefault
class_pathYesPath to .class file
occurrenceNoWhich occurrence to replace (0-indexed)
session_idYesSession ID
value_typeNoValue type hint
replacementsYesMap of old_value -> new_value

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Edit' implies mutation, but it doesn't say whether the change is in-place on disk, scoped to an open session, reversible, or what happens if no occurrence matches - all critical for a bytecode-mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, clearly stating verb and target. It is arguably undersized for a 5-parameter mutation tool, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and a nested `replacements` object, the description omits too much: session requirements (session_id is required), failure behavior, and how the bytecode rewrite is applied. The agent cannot call this confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema itself documents all five parameters, establishing a baseline of 3. The description adds nothing about the semantics of `replacements` (old_value -> new_value mapping), `occurrence` indexing behavior, or `value_type` hints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource: 'Edit numeric constants in .class file bytecode.' This is distinguishable in principle from siblings like jar_edit_constant_pool and jar_hex_edit, but the description never explicitly contrasts itself with them, leaving the agent to infer which editing tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as jar_edit_constant_pool or jar_hex_edit. The agent gets no routing help among several overlapping jar-editing tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_edit_constant_poolC

Edit constants by their exact constant pool index.

ParametersJSON Schema
NameRequiredDescriptionDefault
editsYesMap of pool_index -> new_value
class_pathYesPath to .class file
session_idYesSession ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full disclosure burden, and it provides almost none. It says nothing about whether the edit is destructive/irreversible, whether the class must be recompiled or repacked, whether a session must be open, or what happens if an index is invalid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the mechanism front-loaded and no wasted words. It is not padded, though it is thin on the substance an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool with no annotations and no output schema, so the description should cover preconditions and side effects. It omits session requirements, class_path meaning, failure behavior, and post-edit workflow, leaving the definition materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and all three parameters are documented, so the baseline is 3. The phrase 'exact constant pool index' loosely reinforces the key semantics of the edits map, but the schema already states 'Map of pool_index -> new_value', so no real value is added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Edit) and resource (constants) plus the addressing mechanism (exact constant pool index), which distinguishes it from the sibling jar_edit_class_constants. It stops short of explicitly naming that sibling or clarifying the relationship, so 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (an open session, a valid class path), and no routing to the nearest alternative jar_edit_class_constants or jar_inspect_constant_pool. The agent must infer the trigger context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_edit_fileC

Edit a text file within the JAR.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesFile path within JAR
session_idYesSession ID
new_contentYesNew file content

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden but only implies mutation via 'Edit'. It does not say whether new_content overwrites the whole file, whether changes are in-memory until jar_repack, whether the target file must already exist, or what happens to the session's state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the action front-loaded. It is efficient, though its brevity borders on under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-parameter mutation tool with no annotations and no output schema, the description is too thin. It omits the session-lifecycle context, persistence semantics, and the relationship to jar_repack/jar_close that an agent needs before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both the parameter set and their meanings (file_path, session_id, new_content) are documented in the schema. The description adds no extra semantics such as content encoding or path-relative rules, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb (Edit) and a resource (a text file within the JAR), which is enough to understand the basic action. However, it does not distinguish itself from close siblings like jar_add_file_content, jar_hex_edit, or jar_restore_file, and 'edit' is left ambiguous (full replacement vs. patching) despite a new_content parameter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus jar_add_file_content, jar_hex_edit, or jar_restore_file, nor any stated prerequisites such as needing an open session. The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_hex_editC

Write raw bytes at a specific offset in a file.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetYesByte offset
hex_dataYesHex string to write
class_pathYesPath to file
session_idYesSession ID

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not state that this is a destructive/irreversible in-place mutation, whether a prior jar_open session is required, whether the file must subsequently be repacked, or how bounds/offset errors behave. 'Write raw bytes' alone is insufficient for a mutation tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the operation front-loaded and no filler. It is arguably over-compressed for the surrounding complexity, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-required-parameter mutation tool with no annotations and no output schema, the definition leaves major gaps: session lifecycle, relationship to jar_open/repack, and the exact semantics of offset bounds. The one-sentence description is far too thin for this tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the description adds nothing beyond it, so the baseline of 3 applies. One gap remains: the schema labels session_id only as 'Session ID' without explaining that it ties to an opened jar, and the description does not compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource is clear ('write raw bytes at offset in a file'), but the description is silent about the JAR/class-file context implied by the jar_ prefix and the class_path/session_id parameters. It does not distinguish itself from close siblings hex_write, hex_replace, or jar_edit_file, which an agent must disambiguate among.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative-tool guidance is given. With several overlapping write-style siblings (hex_write, hex_replace, jar_edit_file), the agent gets no signal for choosing jar_hex_edit specifically.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_hex_viewB

Read raw bytes from a .class file as hex dump.

ParametersJSON Schema
NameRequiredDescriptionDefault
lengthNoBytes to read (default 256)
offsetNoStart offset (default 0)
class_pathYesPath to file
session_idYesSession ID

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the read-only nature implicitly ('Read') and the output form ('hex dump'), which is useful, but says nothing about size limits, session requirements, or how the dump is structured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It is appropriately sized for a simple read tool, though it is arguably too terse to earn a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with no annotations and no output schema, the description is minimally adequate: it conveys the operation and output form. It omits session/offset semantics and any return-shape detail, which leaves gaps an agent would otherwise need to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so offset, length, class_path, and session_id are already documented in the schema. The description adds no syntax, format, or constraint detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read), resource (raw bytes from a .class file), and output form (hex dump). It is clearer than a bare name, but does not differentiate itself from close siblings like hex_read, jar_read_file, or jar_read_class, which an agent might reasonably confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (e.g. requiring an open session), and no routing to alternatives such as hex_read or jar_read_file. The agent must infer usage entirely on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_inspect_constant_poolC

Dump the constant pool of a .class file.

ParametersJSON Schema
NameRequiredDescriptionDefault
class_pathYesPath to .class file
session_idYesSession ID
filter_typeNoFilter by constant type
filter_valueNoFilter by value

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Dump' weakly implies a read-only operation, but the description never states that it is non-mutating, that it requires a live session (session_id), or what the output looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is exactly as long as needed and wastes nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, session-based inspection tool with no annotations and no output schema, the description omits the session requirement, the meaning/effect of the filters, and the return format. It is too thin to fully guide invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema. The description adds no extra semantics for filter_type or filter_value, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Dump') and resource ('constant pool of a .class file'), which clearly distinguishes it from the jar_edit_constant_pool / jar_edit_class_constants siblings. However, it never names a sibling or clarifies the boundary with other read tools like jar_read_class.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance. The purpose implies inspection, but the agent gets no signal about when to prefer this over jar_read_class, jar_hex_view, or jar_edit_constant_pool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_listC

List contents of an opened JAR file.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID
max_resultsNoMax results (default 100)
filter_patternNoOptional filename filter

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it discloses only that the JAR must be opened. It says nothing about read-only nature, whether large archives are truncated or paginated, how max_results interacts with listing, or what happens with an invalid/closed session_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the core operation is stated immediately. It is arguably too terse rather than too long, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description is the only source of behavioral detail — and it omits the session prerequisite, the return shape (entry names? paths? metadata?), and result-limiting behavior. For a listing tool with three parameters this leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (session_id, max_results, filter_pattern) are already documented in the schema. The description adds no format, syntax, or matching semantics beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('contents of an opened JAR file'), so the agent knows the operation is an enumeration of archive entries. It does not, however, distinguish itself from close siblings like jar_search, jar_read_file, or jar_list_sessions, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'opened JAR file' implies a prerequisite (jar_open must run first) but never states it explicitly, and there is no guidance on when to choose this over jar_search, jar_read_file, or jar_list_sessions. No exclusions or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_list_sessionsB

List all active JAR sessions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It mentions 'active' sessions, which hints at filtering, but does not disclose whether the operation is read-only, whether it has side effects, what constitutes a session, or what the result format looks like. For a zero-annotation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that wastes no words. It is appropriately sized for a simple list operation, though it could benefit from one additional clarifying clause without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 params, no output schema), the description is minimally adequate but leaves key context unstated. It does not explain what a 'session' is in this context, nor does it describe the return value, which is important since no output schema exists. An agent can guess, but the definition could be more helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool accepts no parameters, so the schema provides no parameter details to supplement. Per the rubric, a zero-parameter tool receives a baseline of 4 for parameter semantics. The description adds no parameter-related information, which is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('active JAR sessions'), making the tool's function clear. However, it does not explicitly differentiate from siblings like jar_open, jar_close, or jar_list, leaving the agent to infer that 'sessions' refers to currently open JAR files rather than, say, files inside a JAR.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. There is no mention of prerequisites (e.g., needing an open session), exclusions, or when to prefer jar_open or jar_close. An agent receives only a bare statement of effect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_openC

Open and extract a JAR file for exploration and editing.

ParametersJSON Schema
NameRequiredDescriptionDefault
jar_pathYesPath to the JAR file
session_idNoOptional custom session ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden, but it only restates the action. It does not disclose whether extraction writes to disk, creates a session, modifies the original JAR, requires specific permissions, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no redundant wording. It is appropriately sized for its limited content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that opens/extracts a JAR and accepts a session_id, the description omits session lifecycle context, whether it is an entry point for other jar_* tools, and any side effects. With no annotations and no output schema, this leaves significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents jar_path and session_id. The description adds no parameter-specific meaning, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Open and extract') and resource ('a JAR file'), plus a scope ('for exploration and editing'). It does not, however, distinguish this tool from siblings like jar_list, jar_read_class, or jar_search, so it falls short of the highest tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the tool is used for exploration and editing, but gives no explicit when-to-use, when-not-to-use, prerequisites, or alternatives. With many sibling JAR tools, an agent has no guidance on when this entry point should be chosen over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_preset_applyC

Apply a saved preset of constant edits.

ParametersJSON Schema
NameRequiredDescriptionDefault
reverseNoReverse the edits (default: false)
session_idYesSession ID
preset_nameYesPreset name to apply

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not say whether edits are written to disk, whether they require an open jar session, whether the operation is reversible (the schema's 'reverse' flag hints at it but the description never explains it), or what happens if the preset is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action and resource, with zero filler. It is efficiently written, though its brevity is closer to under-specification than true conciseness for a mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and an involved jar-editing workflow, this is too thin. It omits whether a jar/session must be open, whether a repack step follows, and what the result state is – all things an agent needs before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – session_id, preset_name and reverse are each documented in the schema. The description adds no syntax, format, or scoping detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource: 'Apply a saved preset of constant edits.' An agent can distinguish this from generic jar editing tools, but it never names or contrasts with its obvious siblings jar_preset_save and jar_preset_list, leaving the preset workflow implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (an open session, a previously saved preset), and no mention of the sibling tools that create and enumerate presets. The agent must infer that a preset must already exist via jar_preset_save.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_preset_listB

List all saved edit presets.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden alone. The verb "List" reasonably implies a read-only, side-effect-free enumeration, but the description says nothing about whether presets persist across sessions, where they are stored, or what the returned entries contain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single six-word sentence with the verb front-loaded and zero filler. Nothing could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should ideally tell the agent what a "saved edit preset" is (a named set of pending JAR edits) and roughly what a listing entry looks like, since that context is not available anywhere else. It is adequate for such a simple tool but leaves the agent guessing about the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics for the description to expand on; the baseline for a 0-param tool with 100% schema coverage is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource combination ("List" + "saved edit presets") that an agent can act on directly. It does not, however, distinguish this from its obvious siblings jar_preset_save and jar_preset_apply, which share the preset namespace; naming them would have sharpened selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no stated when/when-not guidance and no reference to the lifecycle the sibling tools imply (jar_preset_save creates presets, jar_preset_apply applies them). The agent must infer that this is the discovery step before applying a preset.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_preset_saveC

Save a named preset of constant edits for reuse.

ParametersJSON Schema
NameRequiredDescriptionDefault
editsYesEdit map
edit_modeNoEdit mode: constants or pool (default: constants)
class_pathYesTarget class path
descriptionYesPreset description
preset_nameYesPreset name

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but only states that it saves a preset. It does not disclose whether saving overwrites an existing preset, where presets are stored, what permissions are needed, or what the operation returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is appropriately sized for the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has five parameters (including a nested object), no output schema, and no annotations, yet the description is only one sentence. It omits important behavioral context such as persistence location, overwrite behavior, and how presets integrate with jar_preset_apply or jar_preset_list, leaving the definition incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all five parameters in detail. The description adds no further parameter meaning beyond the phrase 'constant edits', which loosely relates to the edits parameter but does not clarify the edit_mode or other fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (save) and resource (named preset of constant edits), so an agent can understand the core action. It does not distinguish itself from sibling tools like jar_preset_apply or jar_preset_list, which is why it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The phrase 'for reuse' hints at a use case, but no conditions, prerequisites, or sibling tool comparisons are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_read_classC

Read and decompile a Java .class file.

ParametersJSON Schema
NameRequiredDescriptionDefault
decompileNoDecompile to Java source (default: true)
class_pathYesPath to .class file within JAR
session_idYesSession ID
extra_classpathNoAdditional classpath entries

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states read/decompile but discloses nothing about the required session state, whether the class must already exist in an open JAR, error behavior, or what the decompile default (already in schema) actually yields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though arguably too terse to be maximally useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful tool requiring a session_id plus a class path, with no annotations and no output schema, the description is too thin. It omits session context, the relationship to jar_open/jar_list, and any notion of what decompilation returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the decompile default. The description adds no syntax or format meaning beyond the schema, which is the expected baseline of 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('read and decompile') and resource (Java .class file), so an agent knows exactly what it operates on. However, it does not distinguish itself from siblings like decompile_type, decompile_method, or inspect_type, leaving the agent to guess which class-level vs type-level tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use context, no prerequisites (e.g. an open JAR session), and no routing guidance relative to the many decompile/inspect siblings. Usage is only implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_read_fileC

Read a text file from the JAR.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesFile path within JAR
session_idYesSession ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It implies a read-only operation and a text-only scope, but never states whether a session must be open, what happens on binary input, error behavior, or size limits for a file-read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action and target front-loaded and zero waste. It is efficient, though its brevity borders on under-specification rather than pure conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool with no output schema, the description covers the core action but omits session requirement, return content, and text-vs-binary handling. It is minimally adequate given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both session_id and file_path are already documented in the schema. The description adds nothing about the file_path format or session context beyond the schema's own text, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('a text file from the JAR'), and the 'text file' qualifier implicitly separates it from jar_read_class and jar_hex_view. However it never names an alternative, so the differentiation from siblings is inferred rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives. The agent can only infer usage from the name and the 'text file' hint; there is no routing information to distinguish this from jar_read_class, jar_hex_view, or jar_search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_repackC

Repack the modified JAR file with all changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID
output_pathNoOptional output path (defaults to original)

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a write/overwrite operation ('Repack... with all changes') but never states what gets destroyed, whether the original JAR is overwritten by default, whether a session must be open, or what happens if no edits were made.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the verb and resource, with no wasted words. It is efficient, though its brevity borders on under-specification for a mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write operation with no annotations and no output schema, the description is too thin: it omits the session requirement, the default overwrite behavior, and the relationship to sibling edit/compile tools. An agent would need to guess at important operational details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both session_id and output_path are already documented in the schema. The description adds no parameter meaning beyond the structured fields, which matches the baseline of 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description has a specific verb ('Repack') and resource ('the modified JAR file'), so the general action is clear. However, it does not distinguish this tool from close siblings like jar_compile_java, jar_close, or jar_restore_file, leaving ambiguity about which repackaging operation applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives such as jar_compile_java or jar_close, nor any stated precondition (e.g. that this must come after jar_edit_file or jar_compile_java). Usage must be inferred entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_restore_fileB

Restore a file from backup to undo modifications.

ParametersJSON Schema
NameRequiredDescriptionDefault
class_pathYesFile path within JAR
session_idYesSession ID

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says 'Restore a file from backup' but does not disclose whether the operation overwrites the current file, what happens if no backup exists, whether the session must be active, or any permissions required. These are significant gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the action and its purpose without any filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations and output schema, the description is too thin for a mutation tool. It omits key context such as session prerequisites, the nature of the 'backup,' and the effect on the JAR file, leaving the agent to infer behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both parameters (session_id, class_path) fully described in the input schema. The description adds no additional meaning beyond what the schema already provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Restore) and resource (a file) plus the purpose (undo modifications). It implicitly distinguishes itself from siblings like jar_edit_file by promising restoration from a backup, but it does not explicitly name an alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'to undo modifications' implies when to use the tool, but there is no explicit guidance on prerequisites (e.g., an open session) or when to choose this over reverting via other means. Usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_scaffold_modC

Generate a minimal Minecraft mod JAR structure.

ParametersJSON Schema
NameRequiredDescriptionDefault
authorNoAuthor name
loaderNoMod loader: fabric, forge, neoforge (default: fabric)
mod_idYesMod ID (lowercase, no spaces)
mod_nameYesMod display name
mc_versionNoMinecraft version (default: 1.21.5)
output_dirYesOutput directory
descriptionNoMod description

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden, and it says nothing about what files are written, whether an existing directory is overwritten, whether the result is immediately loadable, or what permissions/locations are touched. For a tool whose entire purpose is emitting files to disk, this is a substantive gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is efficiently structured, though the terseness is closer to under-specification than true economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With seven parameters, no annotations, and no output schema, the description should say more: what the 'structure' contains (build files, manifest, source dirs), whether it is Java-compiled, and how it relates to subsequent build/deploy tools. An agent cannot confidently chain this tool from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters (including defaults for loader and mc_version) are already documented in the schema. The description adds no syntax, format, or interaction detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Generate a minimal Minecraft mod JAR structure.' An agent can tell it produces a scaffold, but nothing distinguishes it from the similarly named sibling 'scaffold_mod' (nor from 'build_and_deploy' or 'jar_repack'), so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the near-identical sibling 'scaffold_mod'. The agent must guess whether this is the right tool versus scaffold_mod, gorebox_generate_mod, or generate_plugin.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_search_bytecodeC

Search for numeric values or strings in bytecode.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID
file_filterNoOptional file filter pattern
search_valueYesValue to search for

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden and it says nothing behavioral. It does not state whether the match is exact, regex, or case-sensitive, whether an open session is required, or what happens when nothing matches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler and no redundancy. Its brevity is appropriate to the operation, though there is little structure to evaluate beyond that one line.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and three parameters, the description should explain the session dependency, what a result looks like, and how matching works. It delivers none of these, leaving the agent under-informed for a search tool with a required session_id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so session_id, file_filter, and search_value are already documented in the schema; baseline 3 applies. The description adds only the hint that search_value may hold numeric values, which is marginal beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) plus a concrete target (numeric values or strings in bytecode), which is clearer than a bare 'search'. However it never names or distinguishes itself from close siblings like jar_search, jar_search_opcodes, or search_binary_pattern, so an agent must guess which search surface applies to bytecode values.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no routing to alternatives despite four or more sibling search tools. The agent is left to infer that this is the bytecode-wide value search rather than the opcode or binary-pattern searches.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jar_search_opcodesC

Search for bytecode opcode patterns in .class files.

ParametersJSON Schema
NameRequiredDescriptionDefault
patternYesOpcode pattern to search for
session_idYesSession ID
file_filterNoOptional file filter

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about whether the operation is read-only, whether the search is scoped to the given session's loaded jars, how matching is performed, or whether results are capped. Only the surface intent is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no padding or redundancy. It is efficient, though it is arguably too terse for a tool with several ambiguous siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter search tool with fully described parameters and no output schema, the description is minimally adequate. However, given the dense cluster of near-identical search siblings, the absence of any disambiguation or result-shape hint leaves a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so pattern, session_id, and file_filter are already documented in the schema; the description adds no extra meaning such as pattern syntax, whether file_filter supports globs/regex, or how patterns are matched against opcodes. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: searching bytecode opcode patterns within .class files. It is clear on its own, but it does not distinguish itself from the close sibling jar_search_bytecode (or jar_search), leaving the agent to guess which search variant applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the many overlapping search siblings (jar_search, jar_search_bytecode, pattern_scan, search_code). No prerequisites, exclusions, or alternative routing are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_assetsB

List assets in loaded game with file sizes.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 200)
offsetNoSkip first N results (default 0)
path_filterNoFilter assets by path substring
type_filterNoFilter by asset type

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose a prerequisite ("in loaded game") and one output characteristic (file sizes), but says nothing about pagination behavior, whether the call is side-effect free, performance cost, or result ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the scope and the key output detail front-loaded. No filler or redundancy; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, no-output-schema listing tool, the description is minimal but not misleading. It omits how filtering/pagination interact and what a result entry looks like beyond file size, leaving real gaps despite the fully documented schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit, offset, path_filter, and type_filter are already documented in the schema. The description adds no syntax, format, or interaction detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List assets") and adds scope ("in loaded game") plus an output detail ("with file sizes"). It is clear what the tool does, but it never distinguishes itself from close siblings like search_assets, read_asset, or list_types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus search_assets (filtered search) or list_types/list_exports. The only implicit condition is that a game must be loaded, which is not framed as a usage rule. No exclusions or alternatives are offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_available_toolsA

List all tools available for the currently loaded game engine, grouped by category.

ParametersJSON Schema
NameRequiredDescriptionDefault
engineNoEngine to list tools for (uses current game if omitted)

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implicitly discloses a read-only listing behavior and a grouped-by-category return shape, but says nothing about permissions, what happens when no engine is loaded, or whether results are paginated or cached.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler, front-loading the action and resource before the scope qualifier and the return grouping. Nothing could be cut without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, zero-required-parameter read tool with no output schema, the description is nearly sufficient: it conveys the scope and the category grouping of results. It would be complete with one clause on behavior when no engine is loaded or what the categories represent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single optional 'engine' parameter is already documented as defaulting to the current game. The description echoes the default-to-current-engine behavior but adds no syntax, format, or accepted-value detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all tools') scoped to 'the currently loaded game engine', which an agent can act on. It does not differentiate itself from the many sibling list_* tools (list_types, list_monobehaviours, list_exports) or clarify whether 'tools' means engine tooling versus the MCP server's own toolset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the phrase 'currently loaded game engine' signals the prerequisite that a game/engine must already be loaded. There is no explicit when-to-use, when-not-to-use, or pointer to any alternative sibling such as detect_engine or game_status for checking engine state first.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_exportsB

List all exports in a package without reading full properties.

ParametersJSON Schema
NameRequiredDescriptionDefault
asset_pathYesAsset path

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose one real behavioral trait — that properties are not fully materialized, making this a cheap enumeration — and "List" implies read-only. However, it says nothing about return shape, pagination, or error behavior for a missing asset path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the scope and the lightweight-read caveat both appear immediately. It is efficient, though the trailing clause is slightly compressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, no-output-schema listing tool this is close to sufficient, but it never explains what an "export" is in this domain or what the listing returns, leaving an agent to guess the result structure relative to sibling tools like list_assets and read_asset_export.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the schema text is only the bare label "Asset path", adding no format or resolution semantics. The description's "in a package" gives mild context that the path targets a package, but it does not clarify accepted path forms. Baseline 3 for full coverage is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("List") and resource ("exports in a package") and hints at the differentiator from read_asset_export by noting it avoids reading full properties. The only wobble is calling the target a "package" while the parameter is named asset_path, but the intent is still unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase "without reading full properties" implies this is the lightweight alternative to a full export read, which is useful routing context, but no explicit when-to-use, when-not, or named sibling alternative is given. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_monobehavioursC

List all MonoBehaviour components in assembly.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 200)
filterNoOptional name filter
assemblyYesAssembly key

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses very little: not the return shape, not pagination/limit behavior (only the schema hints at a default), not whether an invalid assembly key errors or returns empty, and not any performance cost of scanning an entire assembly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource front-loaded and zero filler. It is efficient, though its brevity is partly a symptom of under-specification rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple read tool with full parameter coverage, but with no annotations and no output schema the description leaves key gaps: what a result item looks like, how limit interacts with total results, and when to prefer this over the sibling type/asset listers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so assembly, filter, and limit are already documented in the schema itself. The description adds no extra semantics (e.g. the format of an assembly key or how filter matching works), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'List all MonoBehaviour components in assembly.' An agent knows exactly what kind of data comes back. However, it gives no differentiation from closely related siblings such as list_types or list_scriptableobjects, so the agent cannot tell which lister to pick without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no stated preconditions (e.g. that the assembly must already be loaded), and no mention of alternatives like list_types or list_scriptableobjects. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_patchable_methodsC

List all methods in a type that can be Harmony patched.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional method name filter
assemblyYesAssembly key
type_nameYesFull type name
include_inheritedNoInclude inherited methods (default: false)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. "List" implies a read-only operation, but the definition never clarifies what qualifies as "patchable" (e.g., excludes abstract/generic/compiler-generated methods), whether inherited members are scanned, or what happens for an unknown type/assembly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler or redundancy. It is appropriately terse for a simple list operation, though that brevity comes at the cost of the gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a Harmony-patching discovery tool with no annotations and no output schema, the description is too thin: it omits the relationship to generate_harmony_patch/validate_patch_target, the return shape (method signatures? patchability reasons?), and any workflow ordering. An agent can guess intent but not usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the input schema, which sets the baseline of 3. The description only loosely echoes that the scope is a type; it adds no semantics for filter, assembly keys, or include_inherited behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ("List all methods") scoped to a type and qualified by the Harmony-patchable criterion, which is meaningfully narrower than generic listing tools. It does not, however, distinguish itself from close siblings like inspect_type or validate_patch_target, so an agent must infer which listing entry point to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance: nothing says whether this is the discovery step before generate_harmony_patch or an alternative to inspect_type. No prerequisites (assembly must be loaded first?) and no exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_scriptableobjectsC

List all ScriptableObject types in assembly.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional name filter
assemblyYesAssembly key

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says nothing about whether the lookup is recursive, how assemblies are resolved, whether the filter is substring or exact, or what the result set looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the resource scoping front-loaded and no filler. It is appropriately sized, though its brevity contributes to the transparency gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should compensate for the missing behavioral and return-value context but does not. It leaves the agent without guidance on result shape, filtering behavior, or error cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents 'filter' as an optional name filter and 'assembly' as the assembly key. The description only restates the assembly scoping and adds no syntax or semantics beyond it, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (List) plus a scoped resource (ScriptableObject types within an assembly). It implicitly separates itself from the broad list_types sibling by narrowing to ScriptableObject types, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no indication of when to prefer list_types or inspect_type, and no prerequisites stated. The agent must infer the use case purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_typesC

List all types (classes, structs, enums, interfaces) in a loaded assembly.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoFilter by kind: class, struct, enum, interface
limitNoMax results (default 200)
filterNoFilter by name (substring match)
assemblyYesAssembly key (from load_assembly)
base_typeNoFilter by base type name

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it does little beyond restating the purpose. It does not disclose pagination behavior (a limit default of 200 exists in the schema), the return shape, or what happens with an invalid/unloaded assembly key.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler and no repetition. It is efficient, though it is minimal enough that it stops short of the structural richness a listing tool could use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter listing tool with full schema coverage and no output schema, the description is adequate but thin. It correctly notes the loaded-assembly precondition, yet omits return/pagination context and offers no routing against closely related type-inspection siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema. The description adds no syntax, format, or semantics beyond the schema (it merely echoes the kinds that 'kind' already enumerates). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (types), and disambiguates the resource by enumerating kinds (classes, structs, enums, interfaces) plus the scope (a loaded assembly). It is clear what the tool does, but it does not distinguish itself from related siblings like inspect_type, decompile_type, or find_renamed_types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'in a loaded assembly' implies a prerequisite (the assembly must be loaded first), but there is no explicit when-to-use guidance and no alternatives named. An agent gets no signal about when to reach for list_types versus inspect_type or get_inheritance_tree.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_assemblyA

Load a .NET assembly (DLL) for analysis. Must be called before other Unity tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
assembly_pathYesFull path to the .NET DLL file
managed_directoryNoOptional: path to the Managed directory for dependency resolution

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses a stateful ordering constraint (a prerequisite call), but says nothing about what happens on re-load, whether loaded state persists across calls, dependency-failure behavior, or permissions. This is meaningful context but far from complete for a mutation/state-establishing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the ordering constraint. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must carry the behavioral load, and it only covers purpose plus the prerequisite ordering. Error cases, state persistence, and what a successful load returns are left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both assembly_path and managed_directory are documented in the schema itself. The description adds no format, path-resolution, or dependency guidance beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource ('Load a .NET assembly (DLL) for analysis'), which is immediately actionable and distinguishable from read-only siblings like analyze_dll_structure. It does not, however, contrast itself against the other loaders in the sibling list (e.g., load_game, jar_open) to clarify which 'load' applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Must be called before other Unity tools' is an explicit prerequisite that tells the agent when this tool belongs in a workflow. It stops short of stating any exclusion or alternative (e.g., what to do if the assembly is already loaded or if this is the wrong loader for the target engine).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_gameA

Set a game as the active modding target. Auto-detects engine, sets up session state, and returns what tools are available for this game type.

ParametersJSON Schema
NameRequiredDescriptionDefault
game_nameNoOptional human-readable game name (auto-detected from path if omitted)
game_pathYesFull path to the game directory

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to lean on, the description still discloses real behavior: it auto-detects the engine, mutates session state, and returns the available tool set for the game type. That is meaningful side-effect and return-shape disclosure, though it omits error behavior (invalid path, already-loaded game) and whether the session can be reset.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the primary action front-loaded and the two secondary effects (engine detection, session setup) trailing. Nothing is wasted and no filler is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a session-establishing tool with no output schema, the description usefully tells the agent what comes back ('what tools are available for this game type'), which is the key missing piece otherwise. It is slightly thin on ordering relative to siblings and on failure modes, but adequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both game_name and game_path are already documented in the schema, and the description adds no format, path-resolution, or engine-detection detail beyond what is there. Baseline 3 applies when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Set a game as the active modding target', which is clearly distinct from read-only siblings like game_status, detect_engine, or open_game. It stops short of explicitly contrasting itself with those siblings, but an agent can still tell what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: calling this before doing mod work is the natural reading, but the description never says when to use it versus open_game, detect_engine, or game_status, nor whether it must precede the mod_* tools. No when-not or prerequisite guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mod_this_gameA

End-to-end game modding assistant. Detects engine, analyzes game structure, and returns a modding plan with available tools and next steps. Start here when modding a new game.

ParametersJSON Schema
NameRequiredDescriptionDefault
game_pathYesFull path to the game directory
objectiveNoWhat you want to mod (e.g. "infinite health", "speed hack", "item spawner")

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the return shape (a modding plan with available tools and next steps) and implies a read/analysis workflow, but it never states whether anything is written to disk, whether the game must be closed, or how long analysis takes — significant unknowns for a mutation-adjacent tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the tool's identity and the action pipeline before the routing cue. Every sentence contributes, though the opening "End-to-end game modding assistant" is more label than information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by summarizing what is returned (a modding plan plus available tools and next steps). For a two-parameter tool with full schema coverage, the main remaining gap is behavioral (side effects, preconditions), which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters carry their own descriptions in the schema, so the structured data already does the work. The description adds nothing about game_path or objective (e.g. whether objective is optional or how it shapes the plan), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (game modding) and enumerates the concrete actions performed: engine detection, structure analysis, and returning a modding plan. It does not explicitly contrast itself with close siblings such as detect_engine, load_game, or scaffold_mod, so an agent must infer that this is the umbrella entry point rather than the narrower tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Start here when modding a new game" gives a clear triggering condition for when this tool is the right entry point. It stops short of stating exclusions (e.g. don't use for an already-loaded game, use load_game instead) or naming alternatives, so it is context-rich but not fully routing-complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

offset_to_rvaB

Convert a file offset to a Relative Virtual Address (RVA) in a PE file.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetYesFile offset to convert
file_pathYesPath to PE file

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says nothing about failure modes (e.g., offset not covered by any section), whether the result is null/error for unmapped offsets, or whether multiple RVAs can map to one offset — all relevant for a PE mapping conversion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and both resources appear immediately. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter conversion with a fully documented schema, the description is minimally sufficient to call the tool. It is incomplete regarding the inverse sibling, offset format expectations, and error behavior for unmapped offsets.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented there (offset = file offset, file_path = PE file path). The description adds no format or unit detail (hex vs decimal offset, expected file type validation) beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (convert) and both resources (file offset, RVA) and scopes it to PE files, so the operation is unambiguous. However, it does not acknowledge the inverse sibling rva_to_offset, so an agent gets no help distinguishing the two directions beyond reading the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus the inverse rva_to_offset tool, nor any prerequisite or context (e.g., which analysis task needs an RVA). Usage must be inferred entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_gameB

Open a game directory for Unreal asset reading. Scans for .pak/.ucas/.utoc files.

ParametersJSON Schema
NameRequiredDescriptionDefault
aes_keyNoAES decryption key if assets are encrypted
game_pathYesPath to game directory
ue_versionNoUnreal Engine version (auto-detected if omitted)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose a real side effect (scanning for .pak/.ucas/.utoc files), which is useful, but says nothing about whether this opens a stateful session, whether later tools depend on it, or what permissions/state are required. That is a significant gap for a stateful loader tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, the core action front-loaded, and the scan targets appended as supporting detail. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the definition is only partially complete: an agent cannot tell whether opening returns a session handle, whether it must precede read_asset/list_assets, or how errors (missing key, no .pak found) surface. The scan behavior is covered, but session/return context is not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains game_path, aes_key, and ue_version. The description adds no format, default, or interaction detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Open a game directory") and narrows the scope to "Unreal asset reading", with the scan targets (.pak/.ucas/.utoc) making the intent concrete. It is distinguishable from Unity/JAR/PE siblings by the Unreal file formats, but it never names a related sibling like load_game or detect_engine, so differentiation is implied rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use framing, no prerequisites (must an AES key or detected engine exist first?), and no mention of alternatives such as load_game or unpack_game. The only guidance is the implicit signal that opening precedes reading assets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pattern_scanB

IDA-style pattern scan with wildcards. Searches for byte patterns in PE sections.

ParametersJSON Schema
NameRequiredDescriptionDefault
patternYesPattern with wildcards (e.g. "48 8B 05 ?? ?? ?? ?? 48 85 C0")
sectionNoPE section to search (default: .text)
file_pathYesPath to PE file
max_resultsNoMax results (default 10)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read-only search but never states it, and says nothing about permissions, cost, scanning limits, or result shape. 'IDA-style' and the .text default are the only behavioral hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, zero waste, with the core operation front-loaded. Nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should at least clarify read-only nature and how results are returned or limited. The fully-documented schema compensates for parameter gaps, but the tool's behavior and sibling routing remain under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented with examples and defaults. The description only reinforces 'wildcards' and 'PE sections', adding little beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Searches for byte patterns in PE sections') plus the flavor ('IDA-style pattern scan with wildcards'). An agent understands what it does, though it never distinguishes itself from the obvious sibling pattern_scan_all or search_binary_pattern.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no exclusions, and no reference to alternatives. With siblings like pattern_scan_all and search_binary_pattern present, the description gives the agent nothing to route between them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pattern_scan_allB

Scan entire file for a pattern (not limited to a section).

ParametersJSON Schema
NameRequiredDescriptionDefault
patternYesPattern with wildcards
file_pathYesPath to file
max_resultsNoMax results (default 50)

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the scan scope (entire file vs section), which is genuinely useful, but says nothing about return format, result ordering, performance on large files, or the max_results cap behavior beyond what the schema already states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the scope qualifier front-loaded after the core action. Nothing is wasted, though the terseness leaves behavioral gaps that a second clause could have filled.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter, no-output-schema, no-annotation tool, the description is minimally viable but thin. It covers what the tool does and its scope, yet omits return shape and any guidance for the optional max_results parameter, leaving the agent to infer the rest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so file_path, pattern, and max_results are already documented in the schema. The description adds no syntax, format, or wildcard details beyond the schema's 'Pattern with wildcards', so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scan entire file for a pattern') and adds a scope qualifier ('not limited to a section') that implicitly distinguishes it from the sibling pattern_scan. The distinction from the sibling is present but relies on the reader inferring what pattern_scan does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(not limited to a section)' hints at when this tool is preferable over the section-scoped pattern_scan, but never names the alternative or states an explicit condition for choosing one over the other. Usage is implied rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_assetC

Read all exports and properties from an asset.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_depthNoMax property nesting depth (default 4)
asset_pathYesAsset path (e.g. /Game/Maps/MainLevel)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Read' implies read-only, but the description does not state that there are no side effects, nor does it mention authentication needs, size/performance limits, or anything about the returned data beyond 'exports and properties'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It is concise, though its brevity means it leaves behavioral and routing details unaddressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter read tool with no output schema and no annotations, the description states what is read but omits when to prefer it over closely related siblings and gives no sense of the return structure. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with only 2 parameters (asset_path and max_depth), so the schema already documents both fully. The description adds no parameter-level meaning, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (asset) plus the scope ('all exports and properties'), which hints at a difference from read_asset_export. However, it never names or explicitly distinguishes itself from sibling tools like read_asset_export, read_blueprint, or list_assets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and no alternatives. An agent cannot tell from the text whether to call this instead of read_asset_export or read_blueprint for a given asset type.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_asset_exportC

Read a specific named export from an asset.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_depthNoMax depth (default 5)
asset_pathYesAsset path
export_nameYesExport name to read

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. Beyond the verb 'Read' implying a safe, non-destructive operation, it discloses nothing about permissions, return format, or how max_depth affects the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though arguably under-specified rather than optimally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (3 params, no nested objects, no output schema), and the schema fully documents parameters, so the description is adequate. But with no annotations it leaves the read-only/safety profile and depth behavior unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no meaning beyond what the schema already documents for asset_path, export_name, and max_depth.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Read) and resource (a named export from an asset), which is clear enough to distinguish from generic siblings like list_assets or read_asset. However, it doesn't explicitly contrast itself with close siblings such as list_exports or read_asset, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and names no alternatives. An agent must guess whether to call read_asset_export versus read_asset, list_exports, or read_blueprint based on the one-line summary alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_blueprintC

Read Blueprint class definition, components, and default properties.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_depthNoMax depth (default 3)
asset_pathYesBlueprint asset path

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read operation and lists the returned content, but says nothing about permissions, whether the blueprint must be loaded first, side effects, or limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the verb and output content front-loaded and no wasted words. It is appropriately sized for a simple two-parameter read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no output schema, the description conveys what data is returned, which is the important part. However, with no annotations and no usage guidance, it leaves gaps around invocation context and constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both asset_path and max_depth are already documented in the schema. The description adds no syntax, format, or meaning beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (Blueprint class), and enumerates the returned content (components and default properties). It is clear what the tool does, but it does not distinguish itself from siblings like read_asset or find_blueprints_of_type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no prerequisites (e.g., asset must be loaded), and no mention of alternative tools. An agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_datatableB

Read a DataTable asset with all rows and field values.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax rows (default 100)
asset_pathYesDataTable asset path
row_filterNoOptional row name filter

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the return content (all rows and field values), which is useful given there is no output schema, but it omits pagination behavior, permission requirements, and the interplay with the limit parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the resource and scope come first. It is efficient, though very terse, providing no structural cues for the optional filtering parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no annotations and no output schema, the description covers the core read behavior but does not explain default limits, error cases, or filter semantics. It is minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so asset_path, limit, and row_filter are already documented in the schema, establishing a baseline of 3. The description adds no parameter detail beyond what the schema provides, and its 'all rows' phrasing sits in mild tension with the default limit of 100.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (read a DataTable asset) and scopes the operation (all rows and field values). However, it does not distinguish this from nearby siblings such as read_asset, read_asset_export, or get_serialized_fields, so an agent must infer which asset-reading tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives like read_asset or read_blueprint, and no prerequisites or exclusions are given. The agent gets a purpose but no routing guidance among the many asset/inspection siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_playerprefsC

Read Unity PlayerPrefs from Windows registry.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoOptional key filter
company_nameYesCompany name from PlayerSettings
product_nameYesProduct name from PlayerSettings

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the Windows-registry backing store, which is useful, but says nothing about read-only guarantees, behavior when PlayerPrefs are absent, permissions, or return format for a tool that crosses into OS-level state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. It is well structured but so terse that it borders on underspecification rather than being optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description should clarify what is returned (key/value pairs?) and what happens on a miss. None of that is present, leaving meaningful gaps for correctness of invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so company_name, product_name and the optional key filter are already documented in the schema. The description adds nothing beyond that, which matches the baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('Unity PlayerPrefs'), and adds the storage location (Windows registry). No sibling tool touches PlayerPrefs, so no differentiation is needed, though it also doesn't differentiate itself from generic read tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any prerequisites such as needing PlayerSettings values or a valid installed game. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_unity_assets_infoC

List Unity asset files in game data directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
data_directoryYesPath to game _Data directory

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only listing via 'List', but does not state whether it reads file contents or only enumerates, what granularity of output results, permission/installation requirements, or how it relates to reading assets with read_asset. For a tool with zero annotation coverage this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the resource front-loaded. It is efficient, though its brevity comes at the cost of the missing usage and behavioral detail noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter listing tool with no output schema this is minimally adequate, but it omits what an 'asset file' encompasses (e.g. .assets files, bundles) and where the tool fits relative to read_asset/list_assets. An agent has enough to attempt the call but not to choose it confidently among siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is documented as the path to the game _Data directory. The description adds no syntax, format, or validation detail beyond what the schema already states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (List) and resource (Unity asset files) with a scoping qualifier (in game data directory), so the agent can tell what it does. However, it offers no differentiation from near-identical siblings such as list_assets, search_assets, or read_asset, leaving ambiguity about which listing tool is Unity-specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus the many sibling listing/reading tools (list_assets, search_assets, read_asset, decode_assets). No prerequisites, no exclusions, no alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rva_to_offsetB

Convert a Relative Virtual Address (RVA) to a file offset in a PE file.

ParametersJSON Schema
NameRequiredDescriptionDefault
rvaYesRVA to convert
file_pathYesPath to PE file

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states nothing about error handling for invalid or out-of-range RVAs, whether the conversion assumes a mapped/loaded image, or any required permissions. For a no-annotation tool this leaves meaningful behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the operation, input, and output with zero filler. Perfectly sized for a simple conversion utility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward two-parameter deterministic conversion with no output schema, the core intent is fully conveyed. Its only shortfall is the absence of error/edge-case behavior, which matters somewhat since no annotations exist to cover it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (both 'rva' and 'file_path' are documented in the schema), so the schema already does the heavy lifting. The description adds no format or unit details beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Convert'), both the source and target resources ('Relative Virtual Address (RVA)' to 'file offset'), and the domain ('PE file'). It is clear enough to distinguish from most siblings, though it does not explicitly name the inverse tool offset_to_rva in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus the obvious alternative offset_to_rva, nor any prerequisites such as whether the PE file must be loaded first. The usage is only implied by the purpose statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scaffold_modC

Generate a mod project skeleton from templates. Creates project files, plugin class, and build configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault
mod_guidNoUnique mod identifier (e.g. com.author.modname)
mod_nameYesMod name (used for class name and project)
frameworkNoMod framework: bepinex5, melonloader, jar (auto-detected from loaded game if omitted)
output_dirYesWhere to create the mod project
game_managed_dirNoPath to game Managed directory (for .csproj references)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only says files are created. It does not state whether existing files in output_dir are overwritten, whether the directory must be empty or pre-exist, or what happens when framework detection fails, which matters for a write-to-disk tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core action front-loaded and a secondary sentence enumerating outputs. No padding, though the second sentence partly restates the first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, five-parameter tool with no annotations and no output schema, the description omits the behaviors an agent most needs: overwrite/merge semantics, success/failure conditions, and any post-generation steps (e.g. whether compile_plugin must follow). Parameters are covered by the schema, but the behavioral picture is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents mod_guid, framework, and game_managed_dir semantics including auto-detection. The description adds nothing parameter-specific beyond implying templates and build config, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Generate a mod project skeleton') plus the artifacts produced (project files, plugin class, build configuration), which is clearer than a bare name. However, it never distinguishes itself from close siblings like generate_plugin or jar_scaffold_mod, so an agent cannot tell which scaffolder to pick from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no prerequisites (e.g. requiring a loaded game or detected engine), and no mention of the alternative scaffolding tools. The agent must infer the call context entirely from the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_assetsC

Search for assets by name (case-insensitive).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 100)
queryYesSearch query

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully notes the match is case-insensitive, but says nothing about what is returned, pagination/limit behavior, or whether the search covers only names versus content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though the extreme brevity leaves it under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema and no annotations, the description should say something about the result set or what an 'asset' is in this domain. As written, an agent cannot predict the return shape or scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description does clarify that 'query' matches on name case-insensitively, adding meaning beyond the schema's bare 'Search query', but the 'limit' parameter is not addressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (assets) plus the search field (name). It does not, however, distinguish itself from sibling tools like list_assets or search_code, so an agent must infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives such as list_assets or read_asset, and no mention of prerequisites or context. The agent is left to infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_binary_patternC

Search for text patterns in a binary file using stream-based approach.

ParametersJSON Schema
NameRequiredDescriptionDefault
patternsYesText patterns to search for
file_pathYesPath to binary file
max_resultsNoMax results per pattern (default 50)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions a 'stream-based approach' but does not explain what that implies (e.g., memory efficiency, handling of large files), nor does it state return format, error handling, or whether the operation is read-only. For a search tool with no annotations, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently states the tool's action and resource. It avoids unnecessary words, though it could benefit from additional concise details about usage or behavior to be more helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations and output schema, the description is incomplete. It does not explain return values, error conditions, or how the stream-based approach differs from other search tools. For a tool with three parameters and no structured behavioral cues, more completeness is needed to guide correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all three parameters, so the baseline is 3. The description adds nothing beyond the schema—it does not clarify pattern matching semantics (e.g., literal vs. regex), encoding expectations, or how 'stream-based' affects the file_path parameter. Thus, it earns the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Search) and resource (text patterns in a binary file), making the core purpose clear. However, it does not differentiate itself from sibling tools like pattern_scan or pattern_scan_all, which likely share overlapping functionality. Without explicit differentiation, an agent must infer distinctions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as pattern_scan or pattern_scan_all. It lacks any usage context, prerequisites, or exclusions, leaving the agent to guess based on tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codeC

Search across assembly for types, methods, fields, or string literals.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 100)
scopeNoSearch scope: all, types, methods, fields, strings (default: all)
patternYesSearch pattern
assemblyYesAssembly key

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full behavioral burden and largely fails it. It does not say whether pattern is a regex, glob, or literal substring, whether matching is case-sensitive, whether the assembly must first be loaded via load_assembly, or what the results look like. Only the implicit read-only nature of "search" is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding or redundancy. It is efficiently written, though its extreme brevity leaves gaps that a slightly longer description could cheaply close.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations and no output schema, the description is thin. It omits prerequisites (is a loaded assembly required?), pattern syntax, result shape/sorting, and any guidance on the interplay between scope and pattern, leaving an agent with real ambiguity about correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (assembly, pattern, scope, limit) is already documented in the schema, and the description's enumeration of searchable categories duplicates the scope parameter's listed values. Baseline 3 is appropriate since the schema does the heavy lifting and the description adds no syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ("Search") and resource ("assembly") and enumerates what is searchable (types, methods, fields, string literals). That is clear enough for an agent to understand the operation, but it never differentiates itself from siblings such as pattern_scan, jar_search_bytecode, or find_references, so an agent cannot route between them from this text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer for overlapping needs (e.g., pattern_scan vs jar_search). The description offers only what the tool does, not when it is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unpack_gameA

CATALOG-FIRST universal unpacker. Recursively walks an ENTIRE game install, classifies and sha256-hashes EVERY file, and persists the catalog into a SQLite .autopsy.db (game/binary/asset/data_store tables, provenance on every row). Read-only: never writes game files. Containers (.pak/.pck/.bundle/...) are catalogued as single rows, never expanded (per-format extraction is a later slice). Idempotent: re-running on the same game re-opens the DB and upserts, no duplicate rows.

ParametersJSON Schema
NameRequiredDescriptionDefault
db_pathNoOptional output path for the .autopsy.db. Defaults to <game_path>/.autopsy/<name>.autopsy.db.
game_pathYesFull path to the game install directory to catalog.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: read-only ('never writes game files'), idempotent (re-runs re-open the DB and upsert, no duplicates), container handling (catalogued, not expanded), and the persistence artifact produced (.autopsy.db with named tables and provenance). This is substantial behavioral disclosure beyond any structured field.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core identity ('CATALOG-FIRST universal unpacker') and each subsequent sentence adding a distinct behavioral fact (scope, read-only, container handling, idempotency). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and no output schema, the description tells the agent what it produces (the autopsy DB and its tables), so return values need not be explained. It omits any note on cost/failure for walking an entire install, which would be useful for a potentially heavy operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, including db_path's default. The description reinforces the default naming (<game>.autopsy.db) but adds no syntax or format meaning beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (recursively walk, classify, hash, catalog) and resource (entire game install) and immediately clarifies the potentially misleading 'unpacker' name: it does NOT expand containers, it catalogs them as single rows. This distinguishes it from siblings like analyze_godot_pck, list_assets, and read_asset, which handle per-format extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies scope ('CATALOG-FIRST', 'per-format extraction is a later slice'), which hints that other tools handle extraction, but it never names a sibling or gives an explicit when-to-use/when-not condition. An agent must infer the routing rather than being told.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unreal_game_statusC

Check current state of Unreal asset reader.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read-only check but says nothing about side effects, whether calling it loads/initializes the reader, what prerequisites exist, or what the result represents — significant gaps for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no wasted words, but the brevity reflects under-specification rather than tight editing. It is front-loaded but carries almost no information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and no parameters, the description is the only source of information, yet it does not explain what 'state' is reported, what triggers it, or what a caller should do with the result. It is too thin for the agent to use confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there is nothing for the description to clarify; the baseline for a zero-parameter tool applies. No parameter meaning is omitted because none exists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb ('Check') and object ('state of Unreal asset reader') are present, but 'state' is undefined — it could mean loaded/unloaded, ready/error, or current asset. It is not distinguished from the very similar sibling 'game_status', so an agent cannot tell the two apart from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus 'game_status', 'load_game', 'detect_engine', or the other Unreal asset tools. The agent must guess whether this is a prerequisite check or a standalone diagnostic.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_assemblyC

Validate compiled plugin DLL for missing references and conflicts.

ParametersJSON Schema
NameRequiredDescriptionDefault
plugin_pathYesPath to compiled plugin DLL
managed_directoryYesPath to game Managed directory

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it only names the validation scope. It does not say what a result looks like (pass/fail, issue list), whether it throws or returns diagnostics, whether it needs the game to be loaded, or what happens when the Managed directory is wrong — all material for a validation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action and its checks front-loaded and no filler. It is well-sized, though the terseness is partly the same under-specification penalized elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must explain the result contract for a validation operation, and it does not. An agent cannot tell what it gets back, how failures are surfaced, or whether it should run before build_and_deploy — gaps that matter for a two-required-parameter validation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented there ('Path to compiled plugin DLL', 'Path to game Managed directory'). The description only reinforces plugin_path indirectly via 'compiled plugin DLL' and adds nothing about managed_directory, so the schema does the heavy lifting — the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (validate) and a specific resource (compiled plugin DLL), plus the exact failure classes it looks for (missing references and conflicts). That is enough to separate it from neighbors like analyze_dll_structure or diff_assemblies, though it does not explicitly name any sibling it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (e.g., that the plugin must already be compiled via compile_plugin or generated via generate_plugin), and no routing to an alternative when the check passes or fails. Usage is only implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_patch_targetB

Validate that a Harmony patch target method exists with expected signature.

ParametersJSON Schema
NameRequiredDescriptionDefault
assemblyYesAssembly key
type_nameYesTarget type name
method_nameYesTarget method name
expected_paramsNoExpected parameter types (comma-separated)
expected_returnNoExpected return type

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. 'Validate' strongly implies a non-destructive read, but the description does not state what happens on a mismatch (error vs. false), whether it requires the assembly to be loaded first, or anything about how the result is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the core purpose front-loaded and zero filler. It is appropriately sized, though it stops short of exploiting the room for useful detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations and no output schema, the description is incomplete: it never explains the return value (boolean, diagnostic message, or error) or how expected_params/expected_return are matched. The schema covers inputs, but the result contract is left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (including expected_params and expected_return) is already documented in the schema. The description mentions 'expected signature' but adds no format or matching-semantics detail beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: validate that a Harmony patch target method exists with an expected signature. An agent grasps the operation immediately. However, it does not name or distinguish itself from closely related siblings like list_patchable_methods, verify_patches, or validate_assembly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, when-not-to-use, or alternative guidance. The domain (Harmony patching) implies a pre-flight check, but nothing tells the agent whether this belongs before generate_harmony_patch, or how it differs from verify_patches/compare_signatures.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_patchesC

Verify Harmony patch targets in plugin still exist in game assembly.

ParametersJSON Schema
NameRequiredDescriptionDefault
plugin_pathYesPath to compiled plugin DLL
game_assemblyYesPath to game Assembly-CSharp.dll

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It conveys that this is a read-only verification (implicitly safe), but says nothing about what a failed verification yields, whether it aborts or reports, or what the output looks like. For a tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is efficient, though the brevity leaves room that could have been spent on the missing usage and output context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter verification tool with no output schema or annotations, the description is minimally adequate but omits the result semantics (success/failure reporting) and any relationship to sibling verifiers. An agent can invoke it but cannot fully predict its behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (plugin_path and game_assembly) are already fully documented in the schema. The description adds no format or constraint detail beyond the schema, matching the baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Verify) and a concrete resource (Harmony patch targets in a plugin against the game assembly), so the operation is unambiguous. However, it does not differentiate itself from the close sibling validate_patch_target, which an agent would need to distinguish to choose correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'still exist' hints at a post-game-update check, but there is no explicit when-to-use, no prerequisite, and no mention of the near-identical sibling validate_patch_target as an alternative. An agent gets no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 100 tool updatesv0.1.4
    • First observedanalyze_bepinex_log
    • First observedanalyze_dll_structure
    • First observedanalyze_file_format
    • First observedanalyze_godot_pck
    • First observedanalyze_pe_full
    • First observedanalyze_save_format
    • First observedbuild_and_deploy
    • First observedcalculate_checksums
    • First observedcompare_binaries_detailed
    • First observedcompare_signatures
    • First observedcompile_plugin
    • First observeddebug_mod
    • First observeddecode_assets
    • First observeddecompile_method
    • First observeddecompile_type
    • First observeddetect_engine
    • First observeddetect_networking
    • First observeddiff_assemblies
    • First observeddisassemble_function
    • First observeddisassemble_range
    • First observedextract_dll_classes
    • First observedextract_strings
    • First observedextract_strings_advanced
    • First observedfind_blueprints_of_type
    • First observedfind_gameplay_values
    • First observedfind_references
    • First observedfind_renamed_types
    • First observedfind_steam_games
    • First observedgame_status
    • First observedgenerate_harmony_patch
    • First observedgenerate_plugin
    • First observedget_class_hierarchy
    • First observedget_inheritance_tree
    • First observedget_method_il
    • First observedget_serialized_fields
    • First observedgorebox_generate_discovery_mod
    • First observedgorebox_generate_mod
    • First observedgorebox_list_mods
    • First observedgorebox_read_api_dump
    • First observedhex_read
    • First observedhex_replace
    • First observedhex_search
    • First observedhex_write
    • First observedinspect_type
    • First observedjar_add_file
    • First observedjar_add_file_content
    • First observedjar_close
    • First observedjar_compile_java
    • First observedjar_detect_mod_info
    • First observedjar_diff
    • First observedjar_edit_class_constants
    • First observedjar_edit_constant_pool
    • First observedjar_edit_file
    • First observedjar_hex_edit
    • First observedjar_hex_view
    • First observedjar_inspect_constant_pool
    • First observedjar_list
    • First observedjar_list_sessions
    • First observedjar_open
    • First observedjar_preset_apply
    • First observedjar_preset_list
    • First observedjar_preset_save
    • First observedjar_read_class
    • First observedjar_read_file
    • First observedjar_repack
    • First observedjar_restore_file
    • First observedjar_scaffold_mod
    • First observedjar_search
    • First observedjar_search_bytecode
    • First observedjar_search_opcodes
    • First observedlist_assets
    • First observedlist_available_tools
    • First observedlist_exports
    • First observedlist_monobehaviours
    • First observedlist_patchable_methods
    • First observedlist_scriptableobjects
    • First observedlist_types
    • First observedload_assembly
    • First observedload_game
    • First observedmod_this_game
    • First observedoffset_to_rva
    • First observedopen_game
    • First observedpattern_scan
    • First observedpattern_scan_all
    • First observedread_asset
    • First observedread_asset_export
    • First observedread_blueprint
    • First observedread_datatable
    • First observedread_playerprefs
    • First observedread_unity_assets_info
    • First observedrva_to_offset
    • First observedscaffold_mod
    • First observedsearch_assets
    • First observedsearch_binary_pattern
    • First observedsearch_code
    • First observedunpack_game
    • First observedunreal_game_status
    • First observedvalidate_assembly
    • First observedvalidate_patch_target
    • First observedverify_patches

TDQS

C2.8/5.0

Scored across 100 tools

Disambiguation2/5

Many tools have overlapping purposes across search, hex, decompilation, and session management. For example, pattern_scan, pattern_scan_all, search_binary_pattern, hex_search, and jar_search_opcodes are all pattern/byte searches, while load_game, open_game, mod_this_game, and detect_engine all initialize game contexts. This creates a high risk of misselection.

Naming Consistency3/5

Most names use snake_case with readable prefixes like jar_ and gorebox_, but the verb/noun ordering is mixed: pattern_scan and hex_read are noun-first, while build_and_deploy and find_gameplay_values are verb-first. It is readable but not a single predictable convention.

Tool Count1/5

100 tools is an extreme surface for one server, well beyond the 15-tool guideline. The set is bloated with overlapping utilities and multiple engine-specific clusters, making discovery and selection unwieldy.

Completeness4/5

The surface covers a broad game-modding lifecycle: engine detection, Unity/Unreal/Godot/JAR analysis, native disassembly, mod generation, compilation, deployment, verification, and repair. Minor gaps exist, such as Godot mod generation/repacking and explicit mod uninstall, but core workflows are covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to perform low-level Windows process memory research, including process attachment, memory scanning, reading/writing, pointer chasing, remote code execution, and inline hooking via MCP tools and Lua scripting.
    2
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Enables AI agents to programmatically control Cheat Engine for memory scanning, editing, cheat table management, and reverse engineering tasks.
    42
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that exposes dynamic binary instrumentation, memory editing, pointer scanning, and scripting capabilities to AI agents, enabling real-time process inspection and modification.
    MIT