Skip to main content
Glama
nickjoven
by nickjoven

erdgraph

A static graph of Elden Ring's executable, served to agents over MCP. It is the mapping layer for an open-source co-op mod: find the code that enforces session rules, name it with evidence, and keep those names alive across game patches.

It works on any MSVC-built PE64, but every default and example targets eldenring.exe.

What it extracts

Ingest takes about 8 seconds and needs no disassembler install.

Source in the binary

What you get

Exception table (.pdata)

Exact function bounds, with cold fragments linked to their parent

Call targets without unwind data

~18k leaf functions, bounds found by forward decode

MSVC RTTI

~12k classes with demangled template names, hierarchy and offsets, ~10k vtables

Code sweep (capstone)

Calls, tail jumps, IAT calls, and RIP-relative reads, writes and address-takes

Base relocations

Every pointer from data into code: initializer tables, callback tables, vtables

Data sections

ASCII and UTF-16 strings, joined to the functions that use them

Export and import tables

560 named Scaleform and Wwise exports, Steam and Win32 imports

The second executable section is the protection layer. It has no unwind data, so it is not part of the function graph.

Related MCP server: HexWitness

Labels and provenance

Names are claims, so they go through a lifecycle instead of being written as facts.

  1. An agent gathers evidence and calls propose_label with a rationale and a confidence.

  2. A separate pass calls review_label with confirms or refutes.

  3. A game patch arrives: ingest the new exe, then port_labels re-finds each name in the new build.

Each event becomes a ket DAG node. Proposals hang off the build node with a proposes edge, and reviews hang off the proposal with confirms or refutes. Ported labels link back with derives. Every event is also appended to labels.log as JSONL.

Porting tries an AOB signature first, with RIP-relative and branch operands wildcarded. Many functions share a shape, so labels also store a string anchor: the longest string referenced only by that function. Strings survive patches far better than code bytes do. Ported labels come in as ported and still need review.

Harvests

Two structural patterns name about 1,700 functions automatically, as proposals:

  • Lua bindings. Static initialisers bind native handlers to Lua event names such as HostDead or StartClientLeaveAroundHost. These come in at confidence 0.7.

  • Qualified-name strings. Debug macros embed names like CS::FieldArea::IsEnableFastTravel. A string used by exactly one function names it, at confidence 0.55.

docs/coop-map.md applies these to the co-op session rules.

Install

You need uv and your own copy of Elden Ring. uv fetches Python 3.12 and prebuilt wheels, so there is nothing to compile.

uv tool install git+https://github.com/nickjoven/erdgraph
erdgraph ingest                 # finds eldenring.exe in your Steam libraries
claude mcp add --scope user erdgraph -- erdgraph serve

ingest reads Steam's libraryfolders.vdf on Windows, WSL and Linux. If the game lives elsewhere, pass the path to eldenring.exe.

ket is optional. With it on PATH, every build and label gets a content ID and lineage. Without it, everything else works and erdgraph says once that provenance is off. Set ERDGRAPH_KET to use a ket binary under another name.

To work on erdgraph itself:

git clone https://github.com/nickjoven/erdgraph && cd erdgraph
uv sync
uv run erdgraph ingest
claude mcp add --scope user erdgraph -- uv --directory "$PWD" run erdgraph serve

Other commands:

erdgraph builds                 # ingested builds
erdgraph harvest --apply        # bulk proposals; about 2 minutes with ket, mostly ket writes
erdgraph port <old> <new>       # carry labels to a new game build

Tools

Tool

Purpose

builds, overview

What has been ingested, counts, label status

find

Substring search over classes, labels, exports, strings, imports

class_info

Bases, derived classes, vtables with named slots, constructor and destructor candidates

function_info

Bounds, callers, callees, imports, strings, globals, vtable slots, data pointers

disassemble

Annotated listing with names resolved inline

xrefs_to, callgraph

Who references an address, and call neighbourhoods

read_memory

Static bytes with pointer naming

make_signature, scan_signature

Unique AOB patterns for hooks and for patch survival

propose_label, review_label, labels, port_labels

The naming lifecycle

global_info

Writers, readers, and the classes whose vtables the writers take, which identifies singletons

harvest_names

Bulk proposals from Lua bindings and embedded qualified-name strings

sql

Read-only SQL over the whole graph

Most tools accept a name where they take an address: a label, an export, a class name for its primary vftable, Class::vfN for a virtual slot, or sub_XXXXXXXX.

Workspace

$ERDGRAPH_HOME, default ~/.local/share/erdgraph:

.ket/                          shared ket store across builds
labels.log                     append-only label events
builds/<blake3[:16]>/image.bin read-only copy of the analysed exe
builds/<blake3[:16]>/graph.sqlite
builds/<blake3[:16]>/manifest.json   hashes, Steam build id, PDB GUID, ket CIDs

The workspace contains a copy of the game executable. Keep it local. The repository holds no game data, and the schema history is in src/erdgraph/schema.changelog.

Limits

  • Static only. Replication timing, packet ordering and authority rules need a runtime harness on the Windows side, which is the next layer.

  • Indirect calls through registers and virtual dispatch are not resolved to targets. Vtable slots and data pointers cover much of that ground.

  • No decompiler. Disassembly is what agents read today. A Ghidra backend needs JDK 21 and is a planned addition.

  • About 2% of RTTI names use forms the demangler skips, mostly member-function pointers inside std::_Binder. They stay mangled and searchable.

Tests

uv run pytest                                          # unit and synthetic
ERDGRAPH_TEST_EXE=/path/to/eldenring.exe uv run pytest # plus real-binary checks

erdgraph is an independent project. It is not affiliated with or endorsed by FromSoftware or Bandai Namco. It ships no game files: it analyses a copy of the executable you already own, and that copy and everything derived from it stay in your local workspace.

It does static analysis only. It does not touch Easy Anti-Cheat, connect to FromSoftware's servers, or modify the game. Anything built on these findings should run offline, as every Elden Ring mod loader already does.

License

MIT. See LICENSE.

Available Tools

18 tools
buildsA

List ingested builds (newest last) with Steam build id and ket manifest CID.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the ordering ('newest last') and the returned fields, but it does not mention permissions, rate limits, or pagination. Since the tool is a simple read operation and an output schema exists, this partial disclosure is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single, tightly written sentence that front-loads the action and includes only relevant details. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity, zero parameters, and an existing output schema, the description provides everything an agent needs to select and invoke the tool correctly. Return values are covered by the output schema, and no further behavioral context is strictly necessary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics for the description to clarify. The baseline for zero parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('ingested builds'), and it states the ordering and the two returned fields. This distinguishes it clearly from all listed sibling tools, none of which deal with builds.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives, nor any preconditions or exclusions. Usage is only implied by the verb 'List'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

callgraphC

Breadth-first call graph around a function. direction: callees | callers.

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
buildNo
depthNo
directionNocallees
max_nodesNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the traversal strategy (breadth-first) and the two traversal directions, which is useful, but says nothing about cost, size limits despite a max_nodes default of 150, or what a call graph result looks like. Significant gaps for a graph-expansion tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short fragments, front-loaded with the core operation and the key enum-like choice. Efficient, though the telegram style (missing articles/verbs) costs a little readability without adding information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with zero schema description coverage, no annotations, and no output schema, this definition is far too thin. Depth, build, and max_nodes semantics and the notion of traversal limits are all left for the agent to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate and largely does not. It only clarifies the direction values (callees/callers) and implicitly the addr target; build, depth, and max_nodes are entirely undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation (breadth-first call graph) scoped to a function, which is a distinct resource from siblings like xrefs_to or function_info. It stops short of naming those siblings or clarifying how the result differs from a plain cross-reference listing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Only hint at usage is the 'direction: callees | callers' note, which reads as parameter documentation rather than a when-to-use condition. No guidance on when to pick this over xrefs_to, function_info, or disassemble, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

class_infoC

RTTI class: bases (with offsets), derived classes, vtables with labelled slots, and constructor/destructor candidates (functions that take the vtable's address).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
buildNo
slot_limitNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It describes the return payload but does not disclose behavioral traits such as read-only status, whether RTTI data must be present, failure modes, or how the build parameter affects behavior. The output detail is useful but falls well short of full behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence fragment with no filler; it lists the key return contents compactly. It is appropriately sized for a lookup tool, though the fragment format is not ideal for scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, 3 parameters, and 0% schema description coverage, the description only covers the return payload. It omits parameter meanings and usage context, leaving an agent unable to reliably invoke the tool without guessing what build and slot_limit control.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention any of the three parameters (name, build, slot_limit). An agent must infer semantics solely from parameter names, so the description adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific RTTI class resource and enumerates what it returns: bases with offsets, derived classes, vtables with labelled slots, and constructor/destructor candidates. This clearly distinguishes it from sibling tools like function_info and global_info. It lacks an explicit verb and does not name an alternative, so it misses the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no alternatives named, and no prerequisites stated. The description only implies that the tool is for RTTI class queries; it does not tell an agent when to choose it over function_info, global_info, or other siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disassembleB

Annotated disassembly. With whole_function, starts at the containing function's start and includes its chained fragments; otherwise starts exactly at addr.

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
buildNo
max_insnsNo
whole_functionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It does disclose the meaningful scoping behavior of whole_function (chained fragments vs. exact addr), which goes beyond the raw boolean. But it says nothing about read-only nature, output shape, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the core purpose and immediately followed by the conditional behavior. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required. For a 4-parameter tool with zero schema coverage, however, the description leaves addr format, build, and max_insns semantics unaddressed, making it only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description must compensate. It meaningfully explains whole_function's behavior, but leaves addr, build, and max_insns (including the 200 default cap) entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Annotated disassembly") that is distinct from siblings like read_memory or function_info. It doesn't explicitly contrast with those siblings, but the disassembly purpose is clear enough to identify the tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the two modes governed by whole_function (function start + chained fragments vs. exactly at addr), which implies when each mode applies. However, it offers no guidance on when to choose this tool over siblings such as read_memory or function_info.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

findB

Substring search (case-insensitive) across classes, labels, exports, strings and imports.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNo
kindsNoclass,label,export,string,import
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses case-insensitive substring behavior and the entity types searched, but omits read-only nature, permissions, side effects, build scoping, and limit behavior. It adds some useful behavioral context but remains incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every word contributes to stating what the tool searches.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with no annotations and 0% schema description coverage, the description leaves build, limit, kinds formatting, and usage context undocumented for a four-parameter search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only implies what the kinds parameter controls by naming the searchable entity categories, and says nothing about build, limit, or query semantics. Most parameter meaning remains undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: case-insensitive substring search across classes, labels, exports, strings, and imports. It is clear what the tool does, though it does not explicitly differentiate itself from siblings like sql or labels.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as sql, labels, or class_info. The scope implies a general search, but no when/when-not conditions or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

function_infoC

Everything the graph knows about the function containing addr: bounds, fragments, labels, vtable slots, callers, callees, imports called, strings and globals touched.

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
buildNo
limitNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must carry the full behavioral burden. It doesn't disclose whether results are read-only (implied), how 'limit' truncates results, how 'build' affects lookup, or whether large functions cause pagination. Only the return categories are implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the core resource and then the returned categories. Efficient, though the trailing list is long and could be trimmed without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three undocumented parameters, no annotations, and no output schema, the description is too thin. It never says how parameters shape the result or what defaults do, leaving an agent guessing about invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters, and the description never explains 'addr', 'build', or 'limit'. With such low coverage, the description needed to compensate but does not mention any parameter at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (function at addr) and enumerates exactly what it returns: bounds, fragments, labels, vtable slots, callers, callees, imports, strings, globals. Clearly distinct from siblings like class_info, global_info, and disassemble.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as disassemble, xrefs_to, callgraph, or class_info. The description implies a broad 'get everything about a function' use case but does not state conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

global_infoC

What lives in a global: writers, readers, and the RTTI classes whose vtables the writers take, which usually identifies a singleton pointer's type. Also the nearest qualified-name strings the writers reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
buildNo
limitNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It adds useful context by listing returned data categories (writers, readers, RTTI classes, nearest qualified-name strings). However, it does not state whether the operation is read-only, what permissions are needed, or how pagination/errors behave, leaving important behavioral traits undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no obvious waste, front-loading the core resource ('What lives in a global'). The jargon is heavy but appropriate for the domain, and the structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter binary-analysis query with no annotations and no output schema, the description leaves major gaps: parameter meanings, usage relative to siblings, and safety profile. It describes returned data categories but is not complete enough for an agent to invoke the tool reliably without opening other artifacts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the parameters addr, build, or limit. It fails to compensate for the undocumented parameters; an agent cannot learn from the description what addr should contain or how build and limit affect results.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States what the tool returns for a global: writers, readers, RTTI classes, and referenced qualified-name strings. This is a specific resource and output scope that distinguishes it from sibling tools like xrefs_to and class_info. It lacks an explicit verb and does not name alternatives, but the purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no guidance on when to use this tool versus siblings such as xrefs_to, class_info, labels, or find. Usage is only implied by the word 'global'. There are no prerequisites, exclusions, or alternative-selection cues.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

harvest_namesB

Propose names from embedded qualified-name strings (e.g. 'CS::FieldArea::IsEnableFastTravel') referenced by exactly one function. Labels land as 'proposed' at confidence 0.55 for review.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNo
dry_runNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses the effect: labels land as 'proposed' at confidence 0.55 for review. However it says nothing about whether the tool mutates persistent state, what dry_run (default true) actually does, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the primary action and effect. The inline example earns its space by clarifying what a 'qualified-name string' looks like. Minor verbosity only.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter coverage, the description carries the full burden. It adequately covers the outcome (proposed labels at 0.55) but omits parameter meaning and mutation scope, leaving meaningful gaps for an agent to close.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither of the two parameters (build, dry_run) is mentioned in the description. The description fails to compensate for the undocumented parameters, so the agent cannot tell what values belong in either field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Propose names') and resource ('embedded qualified-name strings') with a concrete example, so the agent understands it's a batch harvest, not a single manual proposal. It does not explicitly distinguish itself from siblings like propose_label or review_label, which is the only gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The qualifier 'referenced by exactly one function' implies a usage condition for eligible strings, but there is no explicit statement of when to choose this over propose_label or review_label. Usage is inferred from the description rather than directed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

labelsC

List labels, filtered by status (proposed|confirmed|refuted|ported|superseded) and name substring.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNo
limitNo
queryNo
statusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It implies a read listing operation but does not disclose behavior such as pagination (limit param), whether results are sorted, what happens when filters are omitted, or the return structure. An output schema exists, which helps, but the description does not add context beyond the basic purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-structured sentence that front-loads the core action and filter options. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no schema descriptions and no annotations, the description is too sparse. It leaves key parameters (build, limit) unexplained and provides no operational context. An output schema exists, so return values need not be described, but input semantics and usage conditions are incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only clarifies status values and 'name substring' for the query parameter. The build and limit parameters are entirely undocumented, and there is no explanation of default behaviors (limit=200) or how parameters combine.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (List) and resource (labels), and adds the filter dimensions (status, name substring). It is distinguishable from sibling tools like propose_label and review_label which mutate labels, though it does not explicitly call out that distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like find, port_labels, or review_label. The description only states what the tool does, not the context or conditions for selecting it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

make_signatureB

Shortest unique AOB pattern at addr, with RIP-relative and branch operands wildcarded.

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
buildNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses algorithm traits — the pattern is the shortest possible, must be unique, and RIP-relative/branch operands are wildcarded — which is real context. However it omits prerequisites, determinism, and what happens if no unique pattern exists at addr.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the key qualifiers (shortest, unique, addr) front-loaded and zero filler. Appropriately sized for a simple operation, though the terseness contributes to the missing context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description conveys the core output concept (an AOB pattern), which partially substitutes for the absent output schema, but with no annotations, an undocumented build parameter, and no stated preconditions, an agent still lacks enough to invoke this correctly in all cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies 'addr' as the location at which the pattern is generated, but the 'build' parameter is never mentioned or explained in either the schema or the description, leaving half of the parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('make_signature') and defines exactly what is produced: the shortest unique AOB (byte pattern) at an address, with RIP-relative and branch operands wildcarded. This is precise for a reverse-engineering audience, though it never explicitly distinguishes itself from the sibling scan_signature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus scan_signature, or of prerequisites such as needing a loaded binary/module at the given address. The agent must infer usage context entirely from the one-line purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

overviewC

Counts, sections and label status totals for a build (default: newest).

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the default build selection ('default: newest') but omits read-only safety, authentication requirements, pagination, or return format details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is highly concise, though the fragmentary style could be slightly clearer as a full statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with no annotations or output schema, the description at least names the output categories (counts, sections, label status totals). However, it leaves usage, behavioral safety, and return value details implicit, making it only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that the single 'build' parameter selects a build and defaults to the newest, adding meaningful semantics beyond the bare schema, but it does not explain valid formats or how to obtain build identifiers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool's output content ('Counts, sections and label status totals for a build') with a clear resource and scope. It does not differentiate from sibling tools like builds, global_info, or labels, which could cause confusion about which overview-like tool to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The only hint is the default build behavior, which is not a usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

port_labelsC

Carry proposed/confirmed labels from one build to another via their signatures.

ParametersJSON Schema
NameRequiredDescriptionDefault
dst_buildNo
src_buildYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says labels are carried 'via their signatures,' hinting at a matching mechanism, but doesn't disclose what happens on signature mismatches, whether labels are overwritten, whether this requires authentication, or any rate/scope constraints. For a label-migration tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It's a single efficient sentence with no waste, which is good. But at a mere 10 words for a non-trivial tool with zero schema/annotation support, it's under-specified rather than appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 2 parameters, 0% schema coverage, no annotations, and no output schema, the one-line description leaves critical gaps: parameter formats, matching/overwrite behavior, error handling, and how it differs from the many label-related siblings. It is not complete enough for an agent to call it reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no parameter-level detail at all. It mentions 'from one build to another,' which maps loosely to src_build/dst_build semantics, but doesn't explain the required src_build, the optional dst_build (and its empty-string default), or their formats. With two undocumented parameters and no description compensation, this scores minimally.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('Carry') and resources ('labels', 'builds') and the mechanism ('signatures'), which is more specific than the name alone. However, 'labels' is ambiguous terminology — it's unclear whether these are data labels, function labels, or build artifact labels — and the description doesn't distinguish this tool from siblings like labels, propose_label, or review_label.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. There's no mention of prerequisites (e.g., whether src_build must exist or have signatures), when not to use it, or how it relates to the sibling tools that also deal with labels.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_labelC

Record a proposed name with its evidence. Creates a ket node (edge: proposes) under the build node and stores a signature so the label can be ported to later builds.

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
kindNofunction
nameYes
agentNoclaude
buildNo
rationaleYes
confidenceYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses real behavioral traits beyond the name: it creates a ket node, adds a 'proposes' edge under the build node, and stores a signature for portability to later builds. However, it omits whether the operation is idempotent, whether it requires an existing build, error behavior, or the return value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with no filler. It is appropriately sized for the tool, though the domain jargon means it is terse at the cost of clarity rather than because nothing more needs saying.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with zero schema descriptions and no output schema, the description is incomplete. It fails to explain the meaning of the parameters, the preconditions for the build node, or the resulting data shape, leaving the agent under-equipped to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must compensate and largely does not. It implies that 'rationale' is the evidence and that confidence is stored, but the meaning of addr, kind, agent, and build is entirely undocumented in both structured and unstructured text. Only a weak hint is provided for the required fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and object: "Record a proposed name with its evidence." It is distinguishable from siblings like review_label or make_signature, though no explicit sibling differentiation is provided. The purpose is clear but bound to an undocumented domain vocabulary (ket node, build node).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this vs. alternatives, nor when not to. The relation to review_label (likely the human review step) is left to inference. An agent has no way to know the workflow position of this tool without guessing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_memoryC

Static bytes from the image, with qword interpretation and names for pointer-looking values.

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
sizeNo
buildNo

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It adds some context by noting the bytes are 'static' from the image and that values get qword interpretation and pointer-like names, but it omits error behavior, permissions, and whether unmapped or invalid addresses are handled safely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is short but under-specified for a tool with three parameters and no schema descriptions. It is not front-loaded with a clear action verb, and its brevity leaves critical details unaddressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and 0% parameter description coverage, so the description must be more complete. It does not explain how to call the tool, what addr/size/build mean, or what the return structure contains beyond a vague mention of qword and pointer names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the three parameters, but it mentions none of them. The addr, size, and build parameters receive no semantic detail beyond their names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says it returns static bytes from the image and interprets qword/pointer-looking values, which gives a vague sense of the resource. However, it does not explicitly state that it reads memory at a given address, and it does not differentiate this tool from siblings like disassemble or find.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as disassemble, find, or function_info. The description offers no context about prerequisites, appropriate scenarios, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_labelC

Confirm or refute a label ('confirms' | 'refutes'). Creates a ket node linked to the proposal.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoclaude-verifier
buildNo
verdictYes
evidenceYes
label_idYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses only that a 'ket node' is created and linked to the proposal; it says nothing about required permissions, reversibility, what happens on conflicting verdicts, or whether the label state changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the confirm/refute action and its allowed values front-loaded before the side effect. 'ket node' is unexplained jargon but the overall size and ordering are efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no annotations and no output schema, the description omits the meaning of most parameters, the return shape, and any workflow context. Too thin to call this tool correctly without experimentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate but only clarifies the verdict domain ('confirms' | 'refutes'). The other four parameters (agent, build, evidence, label_id) receive no meaning, format, or example beyond their bare titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (confirm/refute) and resource (label), and the parenthetical enumerates the verdict values. It is reasonably distinguishable from propose_label by implication, but never names that sibling, leaving the create-vs-review relationship to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance. 'the proposal' hints at a two-step propose/review workflow with propose_label, but the description never states that this tool is the follow-up to propose_label or what precondition triggers it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_signatureB

Find an AOB pattern ('48 8B 05 ?? ?? ?? ??') in code sections. Returns up to 16 matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNo
patternYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose two useful traits: the search scope is limited to code sections, and results are capped at 16 matches. However, it says nothing about permissions, performance cost, or whether matching is byte-exact vs wildcard-expanded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences: purpose and example first, then the result cap. No filler, no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be described, and the description correctly avoids that. But for a 2-param tool with 0% schema coverage, leaving 'build' undocumented and giving no usage context leaves meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add meaning for 'pattern' by showing an example with wildcard bytes, but the 'build' parameter is left completely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: finding an AOB pattern in code sections, with a concrete example pattern. It's clear what the tool does, but it does not differentiate itself from the sibling make_signature (which presumably creates signatures rather than scanning for them).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no mention of alternatives like find or make_signature. The agent must infer from the purpose alone when scanning is appropriate versus other search tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sqlB

Read-only SQL over the build graph. Tables: functions, classes, class_bases, vtables, vtable_slots, exports, imports, strings, xrefs(insn_va, func_va, dst, kind), labels, sections, meta.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNo
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It usefully discloses 'read-only', which is the key safety trait, but says nothing about query restrictions, row/limit caps, error behavior, or whether arbitrary SQL is permitted against the graph.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences with zero filler; the crucial 'Read-only' qualifier and the subject are front-loaded before the table inventory. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the table list gives solid orientation. Still, for a generic 3-parameter escape-hatch query tool the description omits build/limit semantics and any guidance on query scope or limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. Listing tables and key columns (e.g. xrefs(insn_va, func_va, dst, kind)) meaningfully aids constructing the query parameter, but the 'build' and 'limit' parameters are never mentioned, leaving two of three parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read-only SQL over the build graph') and enumerates the queryable tables, so an agent knows exactly what capability this exposes. It does not, however, differentiate itself from specialized siblings like find, xrefs_to, or function_info that cover overlapping data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. With siblings such as find, xrefs_to, class_info, and callgraph covering similar territory, the description never explains when to reach for raw SQL versus those targeted tools, leaving that choice entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

xrefs_toC

Code references to an address (function, global, string, vtable, IAT slot).

ParametersJSON Schema
NameRequiredDescriptionDefault
addrYes
buildNo
kindsNo
limitNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no statement of result ordering, default truncation (schema default limit=100 suggests capped output), whether addr is a VA or RVA, or pagination. It is not misleading, but it is a near-total omission for an unannotated read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no waste and the key concept is front-loaded, so it is structurally tight. But the brevity is under-specification rather than efficiency, given four undocumented parameters and no behavioral context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter coverage on a four-parameter tool, the description does not supply enough for an agent to call this correctly. The nature of the return (list of reference sites) and any limits on it are left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters (addr, build, kinds, limit), and the description compensates only weakly: the enumerated kinds (function/global/string/vtable/IAT slot) loosely hint at the 'kinds' filter, but addr format, build semantics, and limit behavior are entirely undocumented. The description leaves the agent guessing on most inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Code references to an address' identifies the resource (cross-references) and adds the referenced-entity kinds (function, global, string, vtable, IAT slot), which is more than a bare restatement of the name. However, there is no verb or direction cue beyond the name itself, and no differentiation from siblings such as callgraph or function_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus alternatives like callgraph, function_info, or find, nor any prerequisite or exclusion guidance. The agent must infer usage purely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 18 tool updatesv0.1.0
    • First observedbuilds
    • First observedcallgraph
    • First observedclass_info
    • First observeddisassemble
    • First observedfind
    • First observedfunction_info
    • First observedglobal_info
    • First observedharvest_names
    • First observedlabels
    • First observedmake_signature
    • First observedoverview
    • First observedport_labels
    • First observedpropose_label
    • First observedread_memory
    • First observedreview_label
    • First observedscan_signature
    • First observedsql
    • First observedxrefs_to

TDQS

B3/5.0

Scored across 18 tools

Disambiguation4/5

Most tools target clearly distinct operations (e.g., disassemble vs. callgraph vs. xrefs_to), but there is minor overlap: find can search labels that labels also lists, and propose_label vs. harvest_names both create proposed labels. Descriptions clarify the boundaries, so confusion is limited.

Naming Consistency3/5

Tool names mix noun-style (builds, overview, find, class_info, sql) with verb_noun-style (port_labels, make_signature, scan_signature, propose_label, review_label, harvest_names). The pattern is not consistent, though all names remain readable and understandable.

Tool Count4/5

18 tools is slightly above the typical 3–15 range, but each tool covers a distinct aspect of reverse-engineering graph analysis (builds, classes, functions, disassembly, xrefs, callgraph, globals, memory, signatures, labels, SQL). The count is reasonable for the domain, though not minimal.

Completeness4/5

The surface covers the core lifecycle: build overview, search, class/function/global inspection, disassembly, xrefs, callgraph, signatures, and label proposal/review/listing/porting. Minor gaps exist, such as no explicit label deletion or update operation and no xrefs_from tool, but these are workable via SQL or new proposals.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    A headless Ghidra server that enables AI agents to perform deep reverse-engineering tasks such as disassembly, decompilation, and patching via the Model Context Protocol. It supports extensive automation of analysis workflows in sandboxed environments through a catalog of over 200 specialized tools.
    180
    GPL 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to query persistent, build-scoped evidence memory for reverse engineering, combining static analysis, runtime captures, and claims with honest uncertainty. Provides read-only access to an evidence graph and MCP prompts for structured investigations.
    56 npm
    1
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Enables interactive Ghidra reverse-engineering inside an AI client by rendering live decompiler, function browser, and call graph views. Users can click symbols to rename them and follow calls to navigate, all without leaving the conversation.
    10
    71 npm
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables developer agents to perform semantic codebase search, dependency and impact analysis, cross-file refactoring, and full-stack API tracing through a unified query DSL over a high-performance graph engine.
    -