Chainsaw MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Chainsaw MCP Servertriage the EVTX logs in evidence and summarize the suspicious detections"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
chainsaw-mcp
An MCP server that wraps Chainsaw, the Windows forensic artefact hunting tool, so an agent can triage EVTX event logs, hunt with Sigma and Chainsaw rules, detect log tampering, and author new detection rules without ever pulling raw logs into its context.
Chainsaw release pinned: v2.16.5. MCP spec: 2026-07-28. Python mcp SDK 2.x.
What you get
Area | Tools |
Scope |
|
Hunt |
|
Process pivots |
|
Anti-forensics / timelines |
|
Results (server-minted handles) |
|
Event grouping / oversized rows |
|
Rules |
|
Optional Jev assessment |
|
Resources: chainsaw://rules/{kind}/{path}, chainsaw://mappings/{name},
chainsaw://results/{handle}, chainsaw://docs/rule-format, chainsaw://config.
Prompts: triage_evtx, pivot_on_indicator, author_rule.
Hunts, searches and dumps write JSONL to CHAINSAW_OUTPUT_DIR and return a
res_<hex16> handle plus an aggregate summary and a small preview. Everything else is
paged or grouped through that handle, which keeps the server stateless in the MCP sense
while letting the agent work on tens of thousands of detections.
Pages and exports enforce encoded-JSON byte budgets with resumable offsets; previews are projected rather than full records. EVTX coverage and gap details are handle-backed. Summaries distinguish rule matches from unique source events. See result delivery and continuation.
The implicit default mapping repairs Security 4688 Image translation to
NewProcessName while preserving Sysmon behavior. It derives a private cached mapping
only for the recognized pinned upstream mapping (SHA-256 checked); explicit/custom mapping
selections remain unchanged. The derived file lives under CHAINSAW_OUTPUT_DIR/.mappings/,
is content-addressed and is never swept with expired results. chainsaw_get_mapping
reports routing preconditions and translations.
Related MCP server: Windows Forensics MCP Server
Quick start
uv sync # creates .venv with mcp, httpx, pyyaml, dev tools
uv run chainsaw-mcp-bootstrap # downloads v2.16.5 bundle, verifies sha256,
# installs vendor/chainsaw/{chainsaw,rules,sigma,mappings}
mkdir -p evidence && cp -r /path/to/evtx evidence/
uv run chainsaw-mcp --check # prints status JSON
uv run chainsaw-mcp # stdio transport
uv run chainsaw-mcp --transport streamable-http --host 127.0.0.1 --port 18098 --path /mcpClaude Code registration (stdio):
claude mcp add chainsaw -- uv --directory /path/to/chainsaw-mcp run chainsaw-mcpConfiguration is by environment variables; see .env.example. Evidence paths passed to
tools must resolve inside CHAINSAW_EVIDENCE_ROOTS (default ./evidence). Nothing under an evidence root is ever written.
Transport security: streamable HTTP always runs with DNS-rebinding protection. The bind
host, localhost and 127.0.0.1 are accepted as Host headers; add more with
MCP_ALLOWED_HOSTS (comma-separated, host or host:*) and browser origins with
MCP_ALLOWED_ORIGINS. The endpoint has no authentication of its own, so bind it to a
loopback or private overlay address and let network reachability be the access control;
see SECURITY.md and docs/deployment.md.
The server makes no outbound network calls except the bootstrap download and, when
explicitly enabled, chainsaw_jev_triage.
Jev result triage
chainsaw_jev_triage classifies and prioritizes selected result rows using
TypeSafe Jev. It returns source row indexes,
probabilities and confidence, flags uncertain assessments, and preserves the evidence.
The feature is disabled by default. Enable it with CHAINSAW_JEV_ENABLED=true and either
CHAINSAW_JEV_BWS_SECRET_ID or an injected TYPESAFE_API_KEY. BWS lookup reads the key
only when the tool is called; credentials never appear in status or result output.
Preview a small page with chainsaw_result_page, then call chainsaw_jev_triage with
the same handle, offset, limit and fields. Selected fields are sent to TypeSafe; omitting
fields sends complete selected rows, including the server-side evidence file path. The
tool makes one request for up to 10 rows,
returns advisory scores, and does not automatically run after hunts. See
Jev setup and workflow for configuration, limits and examples.
Deployment
For a long-lived service, run the streamable HTTP transport as a hardened systemd user unit
bound to a private address (a Tailscale IP is the default). deploy/deploy-node.sh ships
a clean, committed tree to a host named by CHAINSAW_DEPLOY_NODE and
CHAINSAW_DEPLOY_USER, syncs the venv, bootstraps Chainsaw if missing, renders the unit
for that host, health-checks the endpoint and rolls back automatically on failure. The
unit is sandboxed and capped at 4 GiB of memory (MCP_MEMORY_MAX) and 512 tasks. The
manual procedure, the unit's sandboxing and the Jev enablement steps are in
docs/deployment.md.
Layout
src/chainsaw_mcp/ server, tools/, rules catalog, mapping repair, result store + byte budgets, bootstrap, CLI
tests/ unit tests (no binary needed; the workflow partition test needs Node), integration/ (needs vendor + samples), fixtures/ (pinned scenario manifest)
docs/ tool reference (generated), result-handle contract, multi-step workflows, Jev setup, deployment
scripts/ quality-check.sh, integration-check.sh, gen-tool-docs.py, check-tool-docs.py, check-package.py, mcp-health.py (endpoint probe)
deploy/ deploy-node.sh (ship over ssh), install-node.sh (target-side install + rollback), render-unit.sh, reference unit
.agents/skills/ agent skills: chainsaw-triage, chainsaw-rule-authoring
.claude/workflows/ Claude Code workflow scripts (evtx-triage, rule-authoring-loop)
.claude/skills/ Claude Code operator skill (/chainsaw)
skills/research/ portable agent skill (frontmatter + markdown), ready to copy into any skills directory
custom-rules/ analyst-authored Chainsaw rules saved by chainsaw_save_ruleDevelopment
bash scripts/quality-check.sh # the CI unit gate: locked sync, ruff, mypy, pytest with coverage, docs drift, bash -n, build
bash scripts/integration-check.sh # the CI integration gate: bootstraps Chainsaw, pins evidence/EVTX-ATTACK-SAMPLES, runs -m integration
uv run pytest # unit tests only
uv run ruff check . && uv run ruff format --check .
uv run mypy # src/chainsaw_mcp, check_untyped_defs
uv run python scripts/gen-tool-docs.py > docs/tools.md # after changing a tool signature, docstring, resource or docs/_result-handles.md
uv run python scripts/check-tool-docs.py # the drift check CI runs.github/workflows/quality.yml runs
both scripts on every push and pull request with pinned action revisions. quality-check.sh fails below 70 % branch coverage and when
docs/tools.md drifts from the live tool schema. docs/tools.md is generated and embeds
docs/_result-handles.md, so edit those sources rather than the generated file.
integration-check.sh fails if any integration test is skipped, so a missing binary or
sample tree can never pass silently. tests/test_workflow_partition.py executes the real
evtx-triage.js in Node and is skipped when node is not on PATH.
Sample evidence lives under evidence/ (gitignored), one directory per public set:
Directory | Source | Contents |
|
| 278 attack-technique files; the integration tests pin this set |
|
| 293 attack files grouped by ATT&CK tactic and technique |
| per-OS | benign goodware logs for false-positive and negative testing |
Contributing and security
See CONTRIBUTING.md for the development workflow and SECURITY.md for the threat model and how to report a vulnerability.
Licence
This server is MIT. Chainsaw itself is GPL-3.0 and is downloaded at bootstrap time, not redistributed here. Sigma rules are DRL 1.1.
Available Tools
26 toolschainsaw_analyse_evtxChainsaw: EVTX overviewA
Summarise EVTX channel, provider and event-ID coverage, including missing metadata.
Runs native analysis and a streamed dump pass. The inline per_file list is a bounded preview; page handle for full files, coverage_handle for identity counts, and diagnostics.handle for all runner messages. Missing Channel in ETW is not evidence of corruption, and successful exit does not establish complete parsing.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes | EVTX files or directories under an allowed root. | |
| skip_errors | No | Continue past unreadable files. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give a partial profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=false), and the description adds real behavior: it runs both a native analysis and a streamed dump pass, the per_file list is bounded, and full output arrives through page/coverage/diagnostics handles. The caveats that missing Channel is not corruption and that exit success does not prove complete parsing are genuinely valuable epistemic context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core summary in sentence one, followed by tightly packed supporting facts about passes, handles and caveats. No padding, though the handle enumeration in the middle sentence is dense and slightly list-like.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by naming the per_file preview plus the page, coverage_handle and diagnostics.handle channels. It also warns about the limits of the result (bounded preview, missing Channel not meaning corruption, exit code not proving full parsing). Missing only explicit guidance on when to prefer this over sibling analysis tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (paths, skip_errors), so the schema carries the parameter burden and a baseline 3 is appropriate. The description adds nothing about path constraints or the meaning of skip_errors beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Summarise) and resource (EVTX channel, provider and event-ID coverage), which cleanly separates it from siblings like chainsaw_analyse_shimcache, chainsaw_analyse_srum and chainsaw_analyse_gaps. An agent can tell what output domain this covers without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the word 'overview' — the description never states when to reach for this versus chainsaw_hunt, chainsaw_analyse_gaps or chainsaw_dump. It does clarify that the inline list is a bounded preview and that full data is fetched via handles, which is useful but not a when-to-use rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_analyse_gapsChainsaw: log gap analysisA
Detect chronological or EventRecordID discontinuities for investigation.
Curated or filtered samples naturally contain gaps; these are not proof of deletion. Correlate with acquisition history and Security 1102 / System 104.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum gaps to return. | |
| paths | Yes | EVTX files or directories under an allowed root. | |
| to_time | No | Analyse up to YYYY-MM-DDTHH:MM:SS. | |
| timezone | No | IANA timezone for output. | |
| from_time | No | Analyse from YYYY-MM-DDTHH:MM:SS. | |
| time_gaps | No | Detect chronological gaps. | |
| skip_errors | No | Continue past unreadable files. | |
| record_id_gaps | No | Detect EventRecordID gaps. | |
| min_time_gap_minutes | No | Flag time gaps at least this long. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false, idempotentHint=false). The description adds non-obvious interpretive context beyond that: gaps in curated/filtered samples are expected and are not proof of deletion, and results should be cross-checked against event IDs 1102/104. It does not, however, note that results are persisted (consistent with readOnlyHint=false).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short lines, front-loaded with the core purpose followed by a concise interpretation caveat and correlation hint. No redundant phrasing, though it is terse enough that return-format expectations are unstated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description does not describe the shape of the result (list of gaps, how they are represented). It does cover the meaning and pitfalls of the output well, which partially compensates, but an agent still lacks confirmation of what it will receive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all nine parameters are self-documented (limit, paths, from/to_time, timezone, time_gaps, record_id_gaps, skip_errors, min_time_gap_minutes). The description adds no parameter-level detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Detect) and resource (chronological or EventRecordID discontinuities), which is clearer than a generic 'analyse EVTX'. An agent can distinguish it from chainsaw_analyse_evtx or chainsaw_hunt, though it never explicitly names a sibling for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for investigation' plus 'Correlate with acquisition history and Security 1102 / System 104' gives implied context for when the output matters. However, there is no explicit statement of when to prefer this over the sibling analysis tools (analyse_evtx, hunt), leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_analyse_shimcacheChainsaw: shimcache timelineA
Build an execution timeline from the shimcache with optional Amcache enrichment.
Wraps chainsaw analyse shimcache. Output is CSV stored under a result handle;
the preview holds the first rows. Use chainsaw_result_export for the full file.
| Name | Required | Description | Default |
|---|---|---|---|
| preview | No | Rows to include inline. | |
| patterns | No | Regexes identifying entries whose shimcache timestamp equals their insertion time (e.g. known-good installer paths). Default: bundled list if present, otherwise none. | |
| system_hive | Yes | SYSTEM registry hive under an allowed evidence root. | |
| amcache_hive | No | Optional Amcache.hve for timeline enrichment. | |
| timestamp_pairs | No | Enable near-timestamp pair detection with Amcache. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false but do not explain why a read-style 'analyse' is not read-only; the description supplies that answer by disclosing that output is a CSV persisted under a result handle with only a preview returned inline. That is meaningful context beyond the annotations, though it says nothing about permissions or runtime cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the implementation, then the output/export contract. No filler and nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return contract (CSV under a result handle, preview inline) and points to the export tool for full retrieval. It is nearly complete for a 5-parameter forensics tool; only the when-to-use selection against sibling analyses is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already carries the parameter burden, and the description adds no syntax or format detail for system_hive, patterns, timestamp_pairs, or preview. The phrase 'the preview holds the first rows' loosely echoes the preview parameter but adds no new semantics, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Build an execution timeline from the shimcache') and names the optional enrichment source (Amcache). The artifact focus distinguishes it from the sibling analyse tools (evtx, srum, gaps) without needing the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent forward ('Use chainsaw_result_export for the full file') which is real guidance, but it never states when to choose shimcache analysis over sibling analyses such as chainsaw_analyse_srum or chainsaw_analyse_evtx, nor any prerequisite conditions. Usage is implied by the tool name rather than explained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_analyse_srumChainsaw: SRUM analysisA
Analyse the System Resource Usage Monitor (SRUM) database.
Wraps chainsaw analyse srum. Output is stored under a result handle; use
chainsaw_result_export for the complete text.
| Name | Required | Description | Default |
|---|---|---|---|
| srum_db | Yes | SRUDB.dat under an allowed evidence root. | |
| stats_only | No | Only report table statistics, not the full extraction. | |
| preview_bytes | No | Bytes of output to include inline. | |
| software_hive | Yes | SOFTWARE registry hive under an allowed evidence root. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnly=false, idempotent=false, destructive=false, but the description explains WHY this non-read-only analysis tool mutates state: output is stored under a result handle. It also clarifies inline output is a preview. That is genuine context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler, and the core purpose is front-loaded before the wrapping detail and the output-routing note. Slightly more could be said about result-handle mechanics, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining that output lands in a result handle and must be retrieved via chainsaw_result_export. Required inputs and their constraints are covered by the schema. Reasonably complete, though it omits any note on large-database runtime or evidence-root requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (srum_db, software_hive, stats_only, preview_bytes) are already documented in the schema. The description adds nothing about parameter semantics; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Analyse the System Resource Usage Monitor (SRUM) database') and names the underlying command. An agent can distinguish it from other analyse_* siblings by resource, though the description never explicitly contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a downstream routing hint ('use chainsaw_result_export for the complete text'), which is useful. However, it gives no guidance on when to choose SRUM analysis over the other analysis tools, and no prerequisites beyond what the schema enforces.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_delete_ruleChainsaw: delete custom ruleADestructiveIdempotent
Delete a custom rule. Bundled Chainsaw and Sigma rules cannot be deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Custom rule file name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds a real operational constraint (bundled Chainsaw/Sigma rules are protected) that prevents a predictable failure, though it does not state permanence or the error behavior when the name is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the action is front-loaded and the constraint follows immediately. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive delete with full annotation coverage and no output schema, the definition is essentially complete. The only minor gap is what a successful deletion returns and whether it is permanent, which the agent could reasonably infer from the destructive/idempotent hints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, so the schema already documents it as the custom rule file name. The description adds no syntax, format, or matching semantics (e.g., name vs. file path), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (custom rule), clearly separating it from the other rule-related siblings (get_rule, save_rule, search_rules) and from result deletion. It stops short of any sibling-by-name routing, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit negative condition: bundled Chainsaw and Sigma rules cannot be deleted, which tells the agent when this tool will fail. There is no positive when-to-use framing beyond the obvious, so it falls short of the 5-level 'explicit when/when-not/alternatives' standard.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_dumpChainsaw: dump recordsA
Convert every record in the evidence to JSON under a result handle.
Wraps chainsaw dump. Use it when a hunt or search has narrowed the scope to one
or two files and you need the complete record stream; page through the handle.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes | Evidence files or directories under an allowed root. | |
| fields | No | Dotted paths or shorthand names to project in the inline preview; see chainsaw_result_fields. Omit for the standard shorthand columns. | |
| preview | No | Records to include inline. | |
| extension | No | Only dump this extension. | |
| skip_errors | No | Continue past unreadable files. | |
| load_unknown | No | Try to parse unidentified files. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag readOnlyHint=false and non-idempotent without explaining why; the description fills that gap by disclosing that output is materialized under a result handle and paginated. This explains the side effect that makes the tool not read-only. It does not detail persistence lifetime or cleanup, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the usage condition immediately following. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still conveys the return model (JSON records under a result handle, paged). Parameters are fully covered by the schema. It leaves the exact paging tool and handle lifecycle implicit, keeping it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all six parameters in detail (fields, preview, extension, skip_errors, load_unknown). The description adds no parameter-level meaning beyond generic paging guidance, matching the baseline 3 for schema-driven tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (convert/dump to JSON) and resource (every record in the evidence), and frames it against the hunt/search workflow. It distinguishes itself contextually as a follow-up step, though it does not explicitly name the sibling tools it is not (e.g. chainsaw_hunt, chainsaw_analyse_evtx).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use condition: after a hunt or search has narrowed scope to one or two files and you need the complete record stream. It also points at the follow-up action (page through the handle). No explicit exclusions or named alternative tools, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_get_mappingChainsaw: read a mappingARead-onlyIdempotent
Describe a Sigma-to-event-log mapping file: groups, filters, field translations.
Mappings decide which Sigma logsources apply to which Windows events. Read this when a Sigma rule you expect is not firing.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Mapping file name. Omit for the effective hunt default. | |
| include_yaml | No | Include the full YAML text. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish that this is a read-only, idempotent operation. The description adds useful domain behavior: mappings determine which Sigma logsources apply to which Windows events, which explains the tool's role beyond simple reading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded: it states what the tool describes, then why mappings matter, then when to use it. Every sentence contributes to selecting or understanding the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only mapping inspection tool with fully documented parameters, the description covers purpose, domain relevance, and usage scenario. It does not detail the return structure, but it does list the main content returned, and no output schema is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters, including the default behavior for name and the purpose of include_yaml. The description does not add parameter-level detail beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Describe') and resource ('Sigma-to-event-log mapping file') and immediately lists the relevant contents: groups, filters, and field translations. This distinguishes it from sibling rule/search tools that operate on rules rather than mapping files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear diagnostic scenario: read this when a Sigma rule you expect is not firing. That is strong when-to-use guidance, though it does not mention alternatives or exclusions, such as searching rules instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_get_ruleChainsaw: read a ruleARead-onlyIdempotent
Read one rule's YAML and parsed summary, or list a rule directory.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | chainsaw, sigma or custom. | |
| path | Yes | Rule path relative to that kind's root, e.g. rules/windows/process_creation/proc_creation_win_lsass_dump.yml or evtx/credential_access. Directories return their children. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is fully covered. The description adds that a directory path returns its children, which is useful behavioral context, but says nothing about error behavior for missing rules or output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence covering both modes with zero waste. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema and only two well-documented params, the description adequately signals both return modes (rule YAML + parsed summary, or child listing). It is nearly complete; only error/edge-case behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters carry examples and the kind enum-like values (chainsaw, sigma, custom), so the schema does the heavy lifting. The description adds no parameter detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (a rule's YAML and parsed summary), plus a secondary listing mode. It is distinguishable from chainsaw_rule_stats and chainsaw_search_rules by implying single-rule retrieval, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The dual mode (read one rule vs. list a directory) is implied, but there is no explicit when-to-use guidance and no alternatives named (e.g., use search_rules instead, or list_evidence for other artifacts). An agent must infer when this beats its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_huntChainsaw: hunt with detection rulesA
Hunt evidence with Sigma and Chainsaw detection rules.
Runs chainsaw hunt over the given paths, stores every detection as JSONL under a
server-minted result handle, and returns aggregate counts (by rule, level, host,
channel, event ID, hour) with a small preview. Prefer this over reading raw events.
| Name | Required | Description | Default |
|---|---|---|---|
| kinds | No | Restrict loaded rules to these kinds: chainsaw, sigma. | |
| paths | Yes | Evidence files or directories, relative to an allowed evidence root or absolute inside one. Directories are searched recursively. | |
| sigma | No | Apply the bundled Sigma rules through the mapping file. | |
| fields | No | Dotted paths or shorthand names to project in the inline preview; see chainsaw_result_fields. Omit for the standard shorthand columns. | |
| levels | No | Restrict rules to these levels: critical, high, medium, low, info. | |
| mapping | No | Mapping file name inside the mappings directory. Default sigma-event-logs-all.yml; sigma-event-logs-legacy.yml for pre-Sysmon logs. | |
| preview | No | Detections to include inline (0-100). | |
| to_time | No | Drop documents newer than this, format YYYY-MM-DDTHH:MM:SS. | |
| statuses | No | Restrict loaded rules to these statuses: stable, experimental. | |
| timezone | No | Render timestamps in this IANA timezone (default UTC). | |
| extension | No | Only load files with this extension, e.g. evtx. | |
| from_time | No | Drop documents older than this, format YYYY-MM-DDTHH:MM:SS. | |
| extra_rules | No | Additional Chainsaw-format rule directories: 'custom' for analyst-saved rules, or 'chainsaw/<subdir>' such as 'chainsaw/evtx/lateral_movement'. | |
| skip_errors | No | Continue past unreadable files instead of failing the hunt. | |
| load_unknown | No | Let chainsaw try to parse files it cannot identify. | |
| chainsaw_rules | No | Apply the bundled Chainsaw rules (rules/ directory). | |
| sigma_collections | No | Sigma collections or sub-trees to load, e.g. ['rules'], ['rules', 'rules-threat-hunting'] or ['rules/windows/process_creation']. Default: ['rules']. Use chainsaw_rule_stats to see what is installed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and the description meaningfully adds to this: it explains that detections are persisted as JSONL under a server-minted result handle and that the response is aggregate counts plus a small preview rather than raw detections. That side-effect (writing results) and the 'result handle' retrieval model are genuinely non-obvious. It stops short of naming the chainsaw_result_* retrieval path beyond a reference to chainsaw_result_fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, with the action and its scope front-loaded and the storage/return behavior immediately after. No filler, and every clause carries information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 17 parameters and no output schema, the description explains what is returned (aggregate counts by rule, level, host, channel, event ID, hour, plus a preview) and where full detections live, which compensates for the missing output schema. Minor gaps remain around how to fetch the full stored result set and rule-loading defaults, but nothing essential to invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; every one of the 17 parameters is documented in the schema itself. The description adds no parameter-level meaning beyond what the schema provides, which is acceptable given the schema's completeness but not above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Hunt evidence with Sigma and Chainsaw detection rules') and follows with the exact underlying command ('Runs `chainsaw hunt` over the given paths'), making the operation unambiguous. It is distinguishable from siblings like chainsaw_search, chainsaw_dump, and the chainsaw_analyse_* family, which do not run detection rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Prefer this over reading raw events' gives a clear directional guideline toward this tool versus raw evidence inspection. However, it never names a concrete sibling alternative (e.g. chainsaw_search or chainsaw_dump) or states when NOT to use it, so the routing guidance is contextual but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_jev_triageChainsaw: assess results with JevA
Send selected result rows to TypeSafe Jev for advisory classification and priority.
EXTERNAL DATA TRANSFER: sends the selected evidence fields to api.typesafe.ai. Requires CHAINSAW_JEV_ENABLED=true and credentials. One bounded API request per call, no retries. Returns source row indexes, probabilities and confidence; invalid answers fail only their source row, with both scores null and a named error; invalid response containers reject the batch. Each assessment lists all uncertain_reasons: insufficient_context, low_classification_confidence, low_priority_confidence or validation_failed. uncertain is true when reasons exist. Scores are model judgments, not confirmed findings. Review the original evidence before acting. Evidence and stored results are unchanged.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum rows sent to Jev. | |
| fields | No | Dotted paths or result-field shorthand to send. Omit to send whole selected rows, including their evidence file path. Use chainsaw_result_page with these fields to preview exactly what would be disclosed. | |
| handle | Yes | Source JSONL result handle to assess. | |
| offset | No | First source row index. | |
| min_confidence | No | Below this confidence, flag uncertainty. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds a lot beyond annotations: names the external destination (api.typesafe.ai), discloses the data-transfer implication, one bounded request per call with no retries, exact failure semantics (invalid answers null both scores with a named error; invalid containers reject the batch), enumerates uncertain_reasons, and cautions that scores are model judgments not confirmed findings. This is unusually thorough for a mutating/external-call tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then the external-transfer warning, then the return/failure semantics. Dense but each sentence carries distinct operational information; the only cost is length from the enumerated uncertainty reasons.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description compensates fully by describing the return shape (source row indexes, probabilities, confidence, uncertain flags and reasons) and per-row failure behavior. An agent has everything needed to invoke and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents limit, fields, handle, offset and min_confidence. The description reinforces the disclosure angle of 'fields' but adds no syntax or format detail beyond the schema, which is the expected baseline-3 outcome when structured data carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb and resource: 'Send selected result rows to TypeSafe Jev for advisory classification and priority.' It also implicitly differentiates from siblings by pointing at chainsaw_result_page for previewing disclosure, so an agent can tell what this tool does versus the result-inspection family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the enabling condition (CHAINSAW_JEV_ENABLED=true plus credentials) and the correct precursor ('Use chainsaw_result_page with these fields to preview exactly what would be disclosed'). No explicit when-not-to-use or comparison against other analysis tools, but the operational context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_lint_rulesChainsaw: lint rulesARead-onlyIdempotent
Validate rules with chainsaw lint and report which files fail to load and why.
Run it after chainsaw_save_rule, or on a Sigma sub-tree to see unsupported modifiers.
| Name | Required | Description | Default |
|---|---|---|---|
| tau | No | Show the compiled tau logic per rule. | |
| kind | No | chainsaw, sigma or custom. | custom |
| path | No | File or directory relative to the kind's root; '.' for all. | . |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds real behavioral value by stating the output nature (files that fail to load and why) and the Sigma unsupported-modifier use case, though it doesn't mention rate limits or output structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action, followed by the usage trigger. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partially compensates by describing what the report contains (failing files and reasons). It stops short of describing return format or how failures are emitted, but it is sufficient for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters (tau, kind, path) are already documented in the schema. The description only indirectly gestures at the Sigma/kind scoping ('Sigma sub-tree') and never explains the tau or kind parameters, so it adds little beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Validate rules with `chainsaw lint`') and specifies the outcome ('report which files fail to load and why'). No sibling offers linting, so an agent can place this tool immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger contexts: after chainsaw_save_rule, or on a Sigma sub-tree to surface unsupported modifiers. It names a sibling for sequencing, but does not state when NOT to use it or the alternatives for pure inspection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_list_evidenceChainsaw: list evidenceARead-onlyIdempotent
List artefact files below an evidence path with sizes and types.
Use it to scope a hunt: pick directories or individual .evtx files instead of hunting an entire mount. Follow next_offset until null for a complete inventory; the evidence tree must remain unchanged between pages.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Directory or file under an allowed evidence root; '.' for the first root. Use chainsaw_status to see the roots. | . |
| limit | No | Maximum files to return. | |
| offset | No | Next offset from the preceding page. | |
| pattern | No | Case-insensitive substring filter on the full path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent/non-destructive profile, so the bar is lower. The description adds genuinely useful behavior beyond that: the pagination contract ('Follow next_offset until null') and a consistency constraint ('the evidence tree must remain unchanged between pages').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action followed by the scoping rationale and pagination rule. No filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema the description carries the burden of explaining returns, and it does state that files come with sizes and types and that pagination continues until null offset. It does not describe the response envelope or the 'next_offset' field's exact shape, leaving a small gap for a paged list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents path, limit, offset, and pattern, making 3 the baseline. The description's 'next_offset until null' hints at the pagination field but uses a name that does not match the 'offset' parameter, adding only marginal clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List artefact files below an evidence path') plus what is returned ('sizes and types'). This is clearly distinguishable from chainsw_hunt/search siblings, which analyze rather than inventory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the intended use: 'Use it to scope a hunt: pick directories or individual .evtx files instead of hunting an entire mount,' which contrasts with the hunt tool by implication. It does not name chainsaw_hunt directly as the alternative, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_process_lineageChainsaw: pivot process lineageA
Pivot process GUIDs in explicitly scoped original evidence, including unmatched events.
Supply files from one investigation (not a shared evidence root) and an exact host. Each hop queries process/parent GUIDs; output is capped and reports incomplete traversal.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes | Explicit evidence files from one investigation. Directories and shared roots are rejected so unrelated scenarios stay separate. | |
| preview | No | Events to include inline (0-100). | |
| computer | Yes | Computer value of the host, compared case-insensitively; other hosts are ignored. | |
| max_hops | No | Parent/child hops to traverse (0-5). | |
| max_events | No | Total events to collect across all hops (1-5000). | |
| max_queries | No | Searches allowed before traversal stops as incomplete (1-256). | |
| process_guid | Yes | Starting Sysmon ProcessGuid, canonical GUID with or without braces. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety hints (non-readOnly, non-destructive, non-open-world), and the description adds genuinely non-obvious behavior: each hop queries process/parent GUIDs, output is capped, and incomplete traversal is reported. The cap/traversal disclosure is exactly the kind of context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short lines, front-loaded with the core action, followed by scope constraints and behavioral caveats. No filler, though 'pivot process GUIDs' leans on insider vocabulary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter read/pivot tool with full schema coverage and no output schema, the description supplies the key behavioral facts (capped output, incomplete-traversal reporting, scope restrictions). It stops short of describing the shape of returned lineage, but that is a minor gap given the constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all seven parameters are already documented with ranges and defaults. The description reinforces the 'one investigation' constraint on paths and 'exact host' on computer but adds no syntax or format detail beyond the schema; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (pivot) and resource (process GUIDs) and scopes it to original evidence, including unmatched events. It is distinguishable from the analyse_* siblings, though a reader must already know Sysmon GUID lineage semantics to feel fully oriented.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a real precondition ('files from one investigation, not a shared evidence root' and 'an exact host'), which is useful guidance. However it never states when to reach for this versus e.g. chainsaw_analyse_evtx or chainsaw_search, so alternative routing is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_chunkChainsaw: read result text chunksBRead-onlyIdempotent
Read UTF-8 text chunks of a file or oversized JSON row. Resume next_byte_offset.
| Name | Required | Description | Default |
|---|---|---|---|
| handle | Yes | Result handle returned by a hunt, search, dump or analysis tool. | |
| max_bytes | No | Content byte budget per chunk (256-131072); UTF-8 characters are never split. | |
| row_offset | No | Source row index of one oversized JSON row to stream (the row_offset an oversized error reports). Omit to stream the stored file itself. | |
| byte_offset | No | Byte position to resume from; pass next_byte_offset from the previous chunk. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering the safety profile. The description adds only that it reads 'UTF-8 text chunks' and supports resuming via next_byte_offset, which is minimal context beyond the schema and annotations. No additional behavioral traits such as rate limits or auth needs are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two short, front-loaded sentences with no wasted words. The core action is stated first, followed by a resume hint, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety and a 100% schema description coverage, the description is adequate but incomplete. It does not explain the workflow context (e.g., obtaining a handle from a hunt/search/dump) or when to choose this tool over other result-reading siblings, leaving an agent to infer its place in the sequence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are fully documented in the schema. The description mentions 'next_byte_offset' (which maps to the byte_offset parameter) but adds no syntax or format details beyond what the schema already provides. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('UTF-8 text chunks of a file or oversized JSON row'), which clearly distinguishes it from most siblings. However, it does not explicitly differentiate from other result-reading tools like chainsaw_result_page or chainsaw_result_events. So it is clear but lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description only says 'Resume next_byte_offset', which hints at resume semantics but provides no explicit guidance on when to use this tool versus alternatives (e.g., chainsaw_result_page). There is no when/when-not context or naming of alternatives, leaving usage largely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_deleteChainsaw: delete a resultADestructiveIdempotent
Delete a stored result and its metadata. Evidence files are never touched.
| Name | Required | Description | Default |
|---|---|---|---|
| handle | Yes | Result handle returned by a hunt, search, dump or analysis tool. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false. The description adds genuinely useful scope beyond that: it says metadata is removed and, importantly, that evidence files survive, which is exactly the reassurance an agent needs before calling a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler; the destructive action comes first and the scoping caveat second.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with full annotation coverage and no output schema, this is nearly complete. It could add whether deletion is recoverable or whether the handle becomes invalid, but the essential scope information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and the schema documents it fully (100% coverage) as the handle returned by a hunt/search/dump/analysis tool. The description adds nothing about the handle, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a stored result and its metadata'), and clarifies the boundary that evidence files are not affected. An agent can distinguish this from the result_* read/export siblings, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: delete a stored result (identified by handle) that you no longer need. There is no explicit when-to-use/when-not guidance, no mention of alternatives (e.g., result_export for preserving output before deletion), and no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_eventsChainsaw: group detections by eventBRead-onlyIdempotent
Page unique source-qualified events, retaining every associated rule match.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Event groups to return. | |
| fields | No | Dotted paths or shorthand names to project inside each group's event. Omit for whole events. | |
| handle | Yes | Result handle returned by a hunt, search, dump or analysis tool. | |
| offset | No | First event-group offset to return (not a detection-row offset). | |
| max_bytes | No | Encoded response byte budget (256-131072). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safe-read profile is covered. The description adds the useful grouping behavior (one row per unique event, with all rule matches retained), but says nothing about ordering, tie-breaking, or how pagination interacts with the retained matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core scope front-loaded and no filler. It is efficient, though arguably compressed to the point of jargon ('source-qualified events') that some agents may find opaque.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With full schema coverage, an existing safety annotation set, and no output schema to explain, the description covers what the tool returns at a high level (event groups with all rule matches). The main missing piece is disambiguation from the other result-viewing siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (limit, fields, handle, offset, max_bytes) is documented in the schema itself, so the baseline is 3. The description adds only the group-vs-row distinction implied by 'Page unique ... events', not new parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Page') and resource ('unique source-qualified events') and adds the key grouping semantics ('retaining every associated rule match'). However, it does not differentiate itself from close siblings like chainsaw_result_page or chainsaw_result_chunk, so an agent cannot tell them apart from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no mention of alternatives, despite the crowded result_* sibling set (result_page, result_chunk, result_fields, result_summary). The agent must infer which result view this tool provides.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_exportChainsaw: export a resultARead-onlyIdempotent
Render a stored result as CSV or JSONL text for reports or other tools.
Large results are capped by rows and encoded response bytes; resume next_offset. CSV pages repeat their header. For stored text files resume next_byte_offset; those UTF-8 chunks may split records and must be concatenated before parsing.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | No | Columns for CSV (dotted paths or shorthand). Default overview set. | |
| format | No | csv, jsonl, or text (an alias of csv). CSV flattens the standard columns of hunts and searches; analysis handles use their sampled top-level keys. | csv |
| handle | Yes | Result handle returned by a hunt, search, dump or analysis tool. | |
| offset | No | ||
| max_rows | No | Row cap for the export. | |
| max_bytes | No | Encoded response byte budget. | |
| byte_offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only/idempotent/non-destructive, and the description adds substantial behavioral detail beyond them: results are row- and byte-capped, CSV pages repeat their header, and text-file UTF-8 chunks may split records and must be concatenated before parsing. These are exactly the operational traits an agent needs and cannot get from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core purpose followed by paging caveats. Every sentence carries useful information about resumption and encoding, though the phrasing is slightly fragmented.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 7 parameters, the description compensates by covering format selection, row/byte caps, and resumption contracts. Missing only explicit guidance on the handle lifecycle and error conditions, which the schema partially covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71% and the description supplements it by explaining the offset/byte_offset resumption semantics and the header-repetition behavior of CSV paging, which the schema titles alone do not convey. It stops short of detailing the fields/format interaction fully, so not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (render/export) and resource (stored result) plus output formats (CSV, JSONL). It does not explicitly name a sibling to distinguish from, but 'for reports or other tools' frames the use case well enough for an agent to identify it as the reporting/serialization path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises when resumption is needed (next_offset for row-capped results, next_byte_offset for stored text files), which implies usage guidance. However it never states when to pick this tool over siblings like chainsaw_result_page or chainsaw_result_chunk, leaving the selection decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_fieldsChainsaw: result fieldsARead-onlyIdempotent
Discover the field paths present in a result so you can project or group on them.
Returns each dotted path with how many sampled rows carry it and an example value.
| Name | Required | Description | Default |
|---|---|---|---|
| handle | Yes | Result handle returned by a hunt, search, dump or analysis tool. | |
| sample | No | Rows to sample. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/closed-world, so the safety profile is covered. The description adds genuine behavioral detail beyond that: it is a sampling operation ('sampled rows') and returns counts plus an example per path, which the agent wouldn't know from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, output shape front-loaded first and the motivating use case second, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param read-only tool with full schema coverage and no output schema, the description is nearly complete: it covers purpose, why to call it, and the shape of the return value. Only the relationship to overlapping siblings (e.g. result_summary) is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters are self-documented, so baseline is 3. The description's phrase 'sampled rows' loosely evokes the sample parameter but adds no syntax, bounds, or default information beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: discover the field paths present in a result. The 'so you can project or group on them' clause scopes the purpose to schema discovery rather than row retrieval, which inherently distinguishes it from the result_events/page/chunk siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (before projecting or grouping) but names no alternatives or exclusions, so the agent must infer that siblings like result_summary or result_page are not the right call here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_listChainsaw: list resultsBRead-onlyIdempotent
List stored result handles, newest first, with tool, inputs and row counts.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Entries per page. | |
| offset | No | First entry to return. | |
| max_bytes | No | Encoded JSON byte budget, including metadata. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint, openWorldHint), so the description need not repeat that. It usefully adds the ordering and the metadata returned per handle, but does not mention pagination behavior or the max_bytes truncation budget.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It immediately conveys the essential action, ordering, and returned metadata.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with rich annotations and full schema coverage, the description is nearly sufficient: it states what is listed, the ordering, and the visible metadata. The only minor gap is that it does not explicitly address pagination or the byte-budget behavior, though the schema covers those details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already explains limit, offset, and max_bytes in enough detail. The description adds no parameter-level meaning beyond what is in the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (List) and resource (stored result handles), plus the ordering (newest first) and included metadata (tool, inputs, row counts). It is clear what the tool returns, though it does not explicitly distinguish itself from siblings like chainsaw_result_page or chainsaw_result_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states what is listed but gives no when-to-use guidance or contrast with alternatives such as chainsaw_result_page, chainsaw_result_summary, or chainsaw_result_fields. An agent must infer the selection context from the tool name and sibling list alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_pageChainsaw: page a resultARead-onlyIdempotent
Read a page of rows from a stored result, optionally projected and filtered.
Rows are chainsaw detections (hunt) or raw documents (search, dump). Iterate with next_offset until it is null.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Rows to return. | |
| where | No | Filters, field path or shorthand -> value; all must match. Text matches as a case-insensitive substring, numbers and booleans exactly (event_id: '4624' matches 4624 only). An aggregate hunt row matches when any of its documents matches. Example: {'level': 'critical', 'computer': 'DC01'}. | |
| fields | No | Dotted paths to project, e.g. ['name', 'level', 'document.data.Event.EventData.CommandLine']. Shorthand names from chainsaw_result_fields also work. Omit for whole rows. | |
| handle | Yes | Result handle returned by a hunt, search, dump or analysis tool. | |
| offset | No | First row index to return. | |
| max_bytes | No | Encoded response byte budget. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring readOnly, idempotent, non-destructive and closed-world behavior, the bar is lower. The description adds real value beyond them: the same tool returns structurally different rows depending on whether the handle came from a hunt, search, or dump, and it exposes the pagination contract (next_offset until null), which is not documented anywhere in the input schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences plus a lead clause, no filler, and the core action is front-loaded. The pagination instruction and the row-type caveat are each independently useful and earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must hint at the return shape; it does so by naming next_offset and the terminal condition. It still doesn't describe the overall response envelope (rows plus metadata), which leaves a small gap for an agent consuming the result, but the critical pagination detail is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (limit, where, fields, handle, offset, max_bytes) is already documented in the schema itself, including the filter matching semantics. The description only alludes to projection and filtering generically ('optionally projected and filtered') without adding syntax or behavior beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read a page of rows from a stored result') plus the projection/filter capability. It clarifies what the rows contain (hunt detections vs raw search/dump documents), which helps orient the agent. It does not, however, distinguish itself from close siblings like chainsaw_result_chunk or chainsaw_result_events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives operative pagination guidance ('Iterate with next_offset until it is null'), which tells the agent how to consume the tool. It stops short of saying when to pick this over chainsaw_result_chunk, chainsaw_result_events, or chainsaw_result_summary, so alternative selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_result_summaryChainsaw: summarise a resultARead-onlyIdempotent
Aggregate a stored result without paging through it.
Use the default overview to understand a hunt, then group by specific event fields (users, processes, source IPs, logon IDs) to build pivots for chainsaw_search.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | Values per grouping. | |
| where | No | Filters applied before counting; same matching rules as chainsaw_result_page's where. | |
| handle | Yes | Result handle returned by a hunt, search, dump or analysis tool. | |
| group_by | No | Field paths or shorthand names to count by, e.g. ['rule', 'computer'] or ['document.data.Event.EventData.TargetUserName']. Omit for the standard overview (rules, levels, hosts, channels, event IDs, hours). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds the "without paging" aggregation trait, but says nothing beyond the schema about cost, limits, or return shape for an aggregation tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, and the core capability (aggregate without paging) is front-loaded before the workflow sentence. Nothing repeats the title or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description names what the default overview yields (rules, levels, hosts, channels, event IDs, hours), so an agent knows the shape of the aggregated response. It does not explain the grouped-mode return format, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so handle, top, where and group_by are already documented, including the default overview contents and the where-matching inheritance from chainsaw_result_page. The description's mention of users/processes/IPs/logon IDs as pivot fields adds only marginal color over the schema's own examples. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ("Aggregate a stored result") with an explicit scope qualifier ("without paging through it") that immediately separates it from the paging siblings like chainsaw_result_page. It also differentiates the two modes it offers: default overview versus group-by pivots.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete workflow: use the default overview to understand a hunt, then group by event fields to build pivots for chainsaw_search. That is clear when-to-use guidance tied to a named sibling. It stops short of saying when NOT to use it (e.g., when you need raw events, use chainsaw_result_events).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_rule_statsChainsaw: rule statisticsARead-onlyIdempotent
Count installed Chainsaw, Sigma and custom rules by level and status, plus mappings.
Also lists the Sigma collections present; pass those names to chainsaw_hunt's sigma_collections when you want threat-hunting or emerging-threat rules included.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuine value beyond that: it discloses the output content (level/status counts, mappings) and that Sigma collection names are listed for reuse in chainsaw_hunt, which is a behavioral fact not in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the core function front-loaded and the follow-up workflow note second. No filler, no restatement of the title or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and annotations covering the safety profile, the description is nearly complete. It tells the agent what the tool produces and how the result feeds chainsaw_hunt; only a brief note on output shape or cost would make it exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify beyond what the empty schema already communicates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: counting installed Chainsaw, Sigma and custom rules broken down by level and status, plus mappings. This is clearly distinct from siblings like chainsaw_search_rules, chainsaw_get_rule, or chainsaw_lint_rules, so an agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit downstream guidance: the listed Sigma collection names should be passed to chainsaw_hunt's sigma_collections for threat-hunting. It gives clear context for what to do with the result, though it does not explicitly contrast with sibling rule-inspection tools such as chainsaw_search_rules.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_save_ruleChainsaw: save custom ruleADestructive
Save an analyst-authored Chainsaw rule into the custom rules directory.
Saved rules are hunted with extra_rules=['custom']. The file is linted before the result is returned; a failing lint still leaves the file in place so it can be fixed with another save using overwrite=true.
| Name | Required | Description | Default |
|---|---|---|---|
| lint | No | Run chainsaw lint on the saved file. | |
| name | Yes | File name without directories, e.g. suspicious_rdp_from_workstation | |
| overwrite | No | Replace an existing custom rule. | |
| yaml_text | Yes | Complete Chainsaw rule YAML with title, group, description, authors, kind, level, status, timestamp, fields and filter. See the chainsaw://docs/rule-format resource. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds non-obvious behavior beyond them: linting runs before the result is returned, and a failing lint leaves the file on disk rather than rolling back, recoverable only via overwrite=true. That partial-failure semantics is exactly the kind of detail an agent cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, with the core action front-loaded and the failure-recovery caveat following immediately. Every sentence carries distinct operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema, the description covers the critical edge case (failed lint still persists the file) and the overwrite remedy, which is what an agent most needs. It stops short of describing the returned result shape or the exact file path, but neither is essential to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents name, yaml_text, lint, and overwrite, including defaults. The description only reiterates overwrite=true as the repair path and confirms lint behavior, adding marginal meaning beyond the schema — baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (save), resource (analyst-authored Chainsaw rule), and destination (custom rules directory), which cleanly separates it from siblings like chainsaw_delete_rule, chainsaw_get_rule, and chainsaw_lint_rules. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the authoring workflow by explaining that saved rules are hunted with extra_rules=['custom'], which is useful downstream context, but never states when to choose this over alternatives or what preconditions apply (e.g. rule must be valid YAML, prior search_rules lookup). Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_searchChainsaw: search eventsA
Search evidence for keywords, regular expressions, or tau field expressions.
Runs chainsaw search; each matching document is stored under a result handle and
summarised by host, channel, event ID and hour. Use tau expressions to pivot on a
specific field (for example a LogonId, process GUID, or user) after a hunt.
| Name | Required | Description | Default |
|---|---|---|---|
| tau | No | Tau field expressions such as 'Event.System.EventID: =4104' or 'Event.EventData.Image: i*mimikatz*'. Combined with AND unless match_any. | |
| paths | Yes | Evidence files or directories under an allowed root. | |
| fields | No | Dotted paths or shorthand names to project in the inline preview; see chainsaw_result_fields. Omit for the standard shorthand columns. | |
| preview | No | Events to include inline. | |
| to_time | No | Drop documents newer than YYYY-MM-DDTHH:MM:SS. | |
| patterns | No | String or regular-expression patterns matched against the whole document. All must match unless match_any is true. | |
| timezone | No | IANA timezone for output. | |
| extension | No | Only load this extension. | |
| from_time | No | Drop documents older than YYYY-MM-DDTHH:MM:SS. | |
| match_any | No | Match a document when any one of the patterns matches (and, separately, any one of the tau expressions) instead of requiring all of them. Patterns and tau expressions are always combined with AND. | |
| ignore_case | No | Case-insensitive pattern matching. | |
| skip_errors | No | Continue past unreadable files. | |
| load_unknown | No | Try to parse unidentified files. | |
| timestamp_field | No | Field holding the timestamp when from_time/to_time are set. Default for EVTX: Event.System.TimeCreated_attributes.SystemTime. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations present, the description adds meaningful behavior beyond them: each matching document is persisted under a result handle and summarised by host, channel, event ID and hour. This side effect is consistent with readOnlyHint=false, and the output shape is disclosed despite no output schema. Auth/rate-limit context is absent, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences: purpose first, then result-handle behavior, then the pivot use case. No filler, nothing repeated, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter tool with no output schema, the description covers purpose, persistence behavior, and the summary shape (host, channel, event ID, hour), which fills the return-value gap. It does not spell out how result handles are consumed or interact with the chainsaw_result_* siblings, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 14 parameters, including tau expression syntax. The description only references the examples already in the schema (LogonId, process GUID, user) and points to chainsaw_result_fields for the fields param, adding little new semantic meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search evidence') plus the three matching modes (keywords, regex, tau expressions), and the second sentence explains what happens to matches. It differentiates from siblings only implicitly via 'after a hunt' rather than naming chainsaw_hunt as the alternative, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use tau expressions to pivot on a specific field ... after a hunt' gives a concrete situation for calling the tool and a worked example (LogonId, process GUID, user). It lacks explicit when-not guidance or a named alternative tool, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_search_rulesChainsaw: search rulesARead-onlyIdempotent
Find detection rules by keyword, ATT&CK tag, level, status or logsource.
Use it to explain a detection name from a hunt, to check coverage for a technique, or to pick a sub-tree for a focused hunt. Returns rule URIs readable as resources.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Substring of an ATT&CK tag, e.g. t1003 or credential | |
| kinds | No | Restrict to rule kinds: chainsaw, sigma, custom. Default all. | |
| level | No | critical, high, medium, low or info. | |
| limit | No | Rules per page. | |
| query | No | Words matched against title, description, tags, path, id. | |
| offset | No | Page start. | |
| status | No | stable, experimental, test or deprecated. | |
| logsource | No | Substring of Sigma logsource values, e.g. process_creation. | |
| path_prefix | No | Rule path prefix, e.g. rules/windows/powershell or evtx/persistence |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so safety is covered. The description adds genuine behavioral value beyond that by disclosing the return format ('Returns rule URIs readable as resources'), telling the agent results can be fed straight into resource reads.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; purpose leads and usage follows. Efficient, though the second sentence is fairly dense with three distinct scenarios.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, all-optional filter tool with a 100%-documented schema, the definition covers purpose, usage scenarios and return format. It leaves minor gaps around how filters combine and pagination expectations, but nothing that would prevent a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already documented in the schema. The description only restates the filter categories (keyword, ATT&CK tag, level, status, logsource) without adding syntax, matching rules, or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Find) and resource (detection rules) plus the filter axes. It distinguishes itself from chainsaw_get_rule (single rule) and chainsaw_hunt only implicitly; no explicit sibling differentiation is offered.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three concrete scenarios (explain a detection name from a hunt, check coverage for a technique, pick a sub-tree for a focused hunt), which is strong usage context. It stops short of naming alternative tools or stating when-not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainsaw_statusChainsaw: server statusARead-onlyIdempotent
Report server configuration, chainsaw version, rule counts and evidence roots.
Call this first to learn which evidence roots and rule sets are available, and whether the chainsaw binary is installed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and non-destructive behavior, so the safety profile is covered. The description adds useful context beyond that (it discloses that the binary may or may not be installed, and that evidence roots/rule sets are discoverable), but says nothing about response shape or failure modes when the binary is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first states what is reported, the second states when to call it. Nothing is repeated and the most decision-relevant instruction ('call this first') is front-loaded in sentence two.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the burden of telling the agent what comes back, and it does list the returned contents (configuration, version, rule counts, evidence roots). A little more detail on the return structure or binary-not-installed behavior would make it fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is no parameter surface for the description to clarify, and nothing is misdescribed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Report server configuration, chainsaw version, rule counts and evidence roots') and enumerates the concrete payload. It implicitly differentiates itself from analysis/hunt siblings by being the configuration-inspection entry point, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear actionable guidance: 'Call this first to learn which evidence roots and rule sets are available, and whether the chainsaw binary is installed.' That establishes the invocation context and prerequisites without naming alternative tools or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
26 tool updates
v0.1.0- First observed
chainsaw_analyse_evtx - First observed
chainsaw_analyse_gaps - First observed
chainsaw_analyse_shimcache - First observed
chainsaw_analyse_srum - First observed
chainsaw_delete_rule - First observed
chainsaw_dump - First observed
chainsaw_get_mapping - First observed
chainsaw_get_rule - First observed
chainsaw_hunt - First observed
chainsaw_jev_triage - First observed
chainsaw_lint_rules - First observed
chainsaw_list_evidence - First observed
chainsaw_process_lineage - First observed
chainsaw_result_chunk - First observed
chainsaw_result_delete - First observed
chainsaw_result_events - First observed
chainsaw_result_export - First observed
chainsaw_result_fields - First observed
chainsaw_result_list - First observed
chainsaw_result_page - First observed
chainsaw_result_summary - First observed
chainsaw_rule_stats - First observed
chainsaw_save_rule - First observed
chainsaw_search - First observed
chainsaw_search_rules - First observed
chainsaw_status
TDQS
Scored across 26 tools
Tools are largely distinguished by artifact type and action: EVTX/gap/shimcache/SRUM analyses, hunt/search/dump, rule management, and result management. Some overlap exists among result_page, result_events, result_export, and result_chunk, but the descriptions clarify their distinct purposes.
All tools use a consistent chainsaw_ prefix and snake_case, making the server easy to navigate. The suffixes mix verb_noun patterns (e.g., chainsaw_delete_rule, chainsaw_analyse_evtx) with noun-first patterns (e.g., chainsaw_rule_stats, chainsaw_result_list), a minor deviation from a strict convention.
At 26 tools, the surface is on the heavy side, with eight result_* tools alone. However, the domain is broad—evidence scoping, hunting, multiple artifact analyses, rule lifecycle, and result paging/export—so the count is borderline rather than clearly excessive.
The set covers evidence listing, hunt/search/dump, targeted artifact analyses, rule reading/searching/saving/deleting/linting/stats, and result listing/paging/chunking/exporting/deleting. Minor gaps include no explicit rule update beyond overwrite-on-save and no analysis tools beyond Chainsaw's supported artifact set.
Maintenance
Related MCP Connectors
Live threat intel for agents: incidents, actors, CVEs with KEV/EPSS, ransomware leak-site victims.
Search log events, investigate anomalies, and manage cases in your Knowledge Grid tenant.
- SuperlogOAuthsh.superlog
Open-source agent that observes and fixes your application. Query logs, traces, metrics, incidents.
Ingest and search LogsLoom logs from coding agents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides programmatic access to ingest and query Windows event logs (especially Sysmon logs), enabling security monitoring, incident response, and log analysis automation.5MIT
- AlicenseBqualityAmaintenanceEnables AI-assisted Windows digital forensics analysis including parsing Windows Event Logs (EVTX), analyzing registry hives (SAM, SYSTEM, SOFTWARE), and remotely collecting artifacts via WinRM with built-in security queries and forensic reference data.4522MIT
- FlicenseBqualityDmaintenanceEnables context-aware EVTX hunting with process lineage tracing and rarity baselining to surface real threats from security logs, transforming raw alerts into actionable kill chain intelligence.31-
- AlicenseAqualityBmaintenanceEnables MCP clients to run Hayabusa detection scans over Windows event log (.evtx) files for forensic analysis and threat hunting.2MIT