Skip to main content
Glama
joblinours

DEFAIR MCP Server

by joblinours

DEFAIR

Digital Forensics & Incident Response platform — MCP-first, containerized, modular.

CI Python 3.13+ License: MIT


What is DEFAIR?

DEFAIR is a reproducible DFIR platform that orchestrates multiple forensic engines, unifies their results, and exposes the investigation to a human or an AI agent via CLI, API, and MCP (Model Context Protocol).

It is not "a giant Docker container with 50 forensic binaries". It is a forensic orchestration platform where external tools are specialized engines managed through a common architecture.

                     Analyst / AI Agent
                            │
                ┌───────────┴───────────┐
                │                       │
               CLI                     MCP
                │                       │
                └───────────┬───────────┘
                            │
                     Service Layer
                            │
                    ┌───────┴────────┐
                    │  Orchestrator  │
                    └───────┬────────┘
                            │
          ┌─────────────────┼─────────────────┐
          │                 │                 │
    Evidence Manager   Tool Registry     Job Engine
          │                 │                 │
          └─────────────────┼─────────────────┘
                            │
                  Normalization Layer
                            │
             ┌──────────────┼───────────────┐
             │              │               │
          Timeline       Findings          IOC
             │              │               │
             └──────────────┼───────────────┘
                            │
                     Reports / Export

Core principles

  1. MCP-first — every capability is exposed via CLI and MCP simultaneously

  2. Read-only on evidence — source files are never modified

  3. Hash & provenance — every result is traceable back to its source

  4. Reproducible — every execution is logged, versioned, and replayable

  5. No shell via MCP — the MCP exposes forensic operations, not arbitrary commands

  6. Offline-first — designed to work without internet access


Related MCP server: Forensic Artifact Investigator MCP Server

Quick start

Installation

# Clone
git clone https://github.com/joblinours/defair.git
cd defair

# Create venv and install (dev includes pytest, ruff, etc.)
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

# Verify installation
defair --version
pytest tests/ -v

CLI usage

# Create a forensic case
defair case create "Incident 2026-09 — Host compromise"

# List cases
defair cases list

# Register evidence (computes SHA-256 automatically)
defair evidence add CASE-2026-001 /evidence/host01.E01 --type disk_image

# List evidence (all or filtered by case)
defair evidence list
defair evidence list --case CASE-2026-001

# Get details
defair case get CASE-2026-001
defair evidence get EVD-001

# Verify evidence integrity (re-hash and compare)
defair evidence verify EVD-001

# Whole investigation in one command (v0.4): collection folders, archives,
# Generaptor / DFIR-ORC, disk images — prepared, then a profile runs in the background
defair run start --case CASE-2026-001 --evidence /evidence/host01.E01 --profile auto
defair run status PRUN-001

Container orchestration

DEFAIR runs forensic tools in isolated Docker containers — one per investigation. Containers are hardened by default: all capabilities dropped, no-new-privileges, no network, read-only root filesystem, CPU/memory/PID limits, running as your host user. Evidence can only be mounted from the directories listed in container.evidence_roots (defair.yaml). Hardening applies to containers created from v0.3.6 on — recreate older ones.

# Create a forensic container linked to a case
defair container create --case CASE-2026-001 --evidence /evidence/host01.E01

# List containers
defair container list

# Interactive analyst shell (CLI only — never exposed through MCP)
defair container shell defair-case-2026-001

# Execute a command inside the container
defair container exec defair-case-2026-001 ls -la /evidence

# Get container details / logs
defair container get defair-case-2026-001
defair container logs defair-case-2026-001

# Stop / remove
defair container stop defair-case-2026-001
defair container remove defair-case-2026-001

MCP usage

DEFAIR exposes a MCP server over stdio — compatible with Claude Desktop, Claude Code, and any MCP client:

# Run the MCP server
defair-mcp

Available MCP tools:

Tool

Description

Case & Evidence

create_case

Create a new investigation case

list_cases

List all forensic cases

get_case

Get case details by ID or case number

register_evidence

Register evidence with SHA-256 hash

list_evidence

List evidence (optionally filtered by case)

get_evidence

Get evidence details by ID or number

verify_evidence

Re-hash evidence and verify integrity

Container

create_container

Create an isolated forensic container

list_containers

List DEFAIR containers

get_container_info

Get container details

start_container

Start a stopped container

stop_container

Stop a running container

remove_container

Remove a container

run_tool

Run a registered tool with validated options (recorded as a ToolRun)

exec_in_container

Arbitrary command — disabled unless mcp.allow_exec: true

container_logs

Get container logs

Discovery & Analysis

discover_evidence

Discover forensic artifacts on mounted evidence

list_tools

List available forensic tools

tools_health

Check tool availability in a container

list_tool_runs

List past analysis runs

list_artifacts

Search artifacts (text, tool, host, user, time range, run) with paging

get_artifact

Deep-inspect one artifact: data, provenance, raw EVTX event / YARA hex, timeline context

get_finding

A finding with every match explained (file, event / offset, matched values, rule file)

show_rule

Content of a YARA / Sigma rule from the verified store

analyze_evtx

Parse Windows Event Logs (EvtxECmd)

analyze_mft

Parse NTFS Master File Table (MFTECmd)

analyze_registry

Parse Windows Registry hives (RECmd)

analyze_prefetch

Parse Prefetch files (PECmd)

analyze_amcache

Parse Amcache.hve (AmcacheParser)

analyze_shimcache

Parse Shimcache (AppCompatCacheParser)

analyze_jumplist

Parse Jump Lists (JLECmd)

analyze_lnk

Parse LNK shortcuts (LECmd)

analyze_recyclebin

Parse Recycle Bin (RBCmd)

analyze_shellbags

Parse ShellBags (SBECmd)

analyze_srum

Parse SRUM database (SrumECmd)

analyze_wintimeline

Parse Windows Timeline (WxTCmd)

analyze_sqlite

Parse SQLite databases (SQLECmd)

Detection & Hunting (v0.3)

hunt_evtx

Run Hayabusa Sigma detection on EVTX

build_timeline

Build unified timeline summary

search_timeline

Search/filter the timeline (artifacts + Plaso / Sleuth Kit events)

list_findings

List investigation findings

search_ioc

Search IOC across all artifacts

Scanning & rules (v0.3.7)

scan_yara

YARA scan of every file (Raijin, pinned rule sets)

scan_sigma

Sigma scan of EVTX / Linux logs (KAPE, Velociraptor, mount)

scan_evidence

YARA + Sigma in a single pass

get_ruleset_info

Pinned rule sources, installation and integrity status

list_rule_conflicts

Duplicate / conflicting rules across sources

Investigations (v0.4)

analyze_evidence

Path → registered, prepared, profile chosen and run (background)

prepare_evidence

Extract / decrypt / carve an evidence (ZIP, Generaptor, DFIR-ORC, disk images)

list_profiles / get_profile

Analysis profiles and their steps

run_profile

Run a profile on an evidence in the background (PRUN-NNN)

get_run_status / list_runs

Follow profile runs (steps, tools, fallbacks, errors)

cancel_run / resume_run

Stop a run / resume it without re-running completed steps

Windows coverage (v0.4.5)

analyze_usn

USN journal ($J) with MFTECmd, parent paths from the $MFT

analyze_ual

User Access Logging (Windows Server) with SumECmd

hunt_chainsaw

Sigma hunt with Chainsaw on the pinned rule store

list_evtx_views / get_evtx_view

Typed EVTX views (logons, services, RDP, PowerShell…)

get_host_profile

Host profile with a source per fact

list_watchlists / search_watchlist

Keyword / IOC watchlists over artifacts, strings, raw files

Supertimeline (v0.5)

build_supertimeline

Plaso / Sleuth Kit job in the worker-plaso image (returns a RUN-NNN)

get_supertimeline_status

Job state; imports the events into the timeline when finished

cancel_supertimeline

Stop a supertimeline job

Normalization & export (v0.3.8)

export_timeline

Export the timeline (Timesketch JSONL, JSONL, CSV)

normalize_replay

Rebuild a case's artifacts from its normalized JSONL

normalize_rerun

Re-normalize a tool run from its raw output

get_normalization_stats

Normalization counters of a run

Docker

# Build
docker compose build

# Run CLI
docker compose run --rm defair defair case create "Docker test"

# Run MCP server
docker compose run --rm mcp

Architecture

DEFAIR follows a triple-interface architecture: CLI, MCP, and API all call the same async Service Layer. No business logic lives in the interface layers.

CLI (click)          MCP (FastMCP)          API (FastAPI — planned)
     │                     │                      │
     │  run_sync()         │  async direct         │  async direct
     └─────────┬───────────┴──────────────────────┘
               │
        Service Layer (async)
               │
     ┌─────────┴──────────┐
     │                    │
  SQLite            Docker SDK
  (aiosqlite)       (container_service)
                          │
              ┌───────────┴───────────┐
              │  DEFAIR Container 1   │  evidence :ro
              │  DEFAIR Container 2   │  workspace :rw
              └───────────────────────┘

Project structure

src/defair/
├── config.py              # YAML + Pydantic configuration
├── logging.py             # Structured logging (structlog)
├── database.py            # SQLite schema & connection management
├── models/                # Pydantic data models (Case, Evidence, ...)
├── services/              # Async service layer (shared by CLI & MCP)
├── cli/                   # Click CLI commands
└── mcp_server/            # FastMCP server & tool definitions

Tech stack

Component

Technology

Language

Python 3.13

CLI

Click + Rich

MCP server

FastMCP 4.x (stdio)

Database

SQLite via aiosqlite

Models

Pydantic v2

Logging

structlog (JSON / console)

Config

YAML + Pydantic

Tests

pytest + pytest-asyncio

Lint

Ruff

Container

Docker + Compose

CI/CD

GitHub Actions


Roadmap

DEFAIR is built MCP-first: every phase delivers the forensic capability and its MCP exposure simultaneously.

The roadmap below also closes the coverage gap with all-in-one DFIR toolboxes such as Hecatrace — without becoming one. Each tool they ship is integrated the DEFAIR way:

  • Wrapped, not exposed — a BaseTool manifest + normalizer, never a raw binary reachable through MCP

  • Structured, not dumped — results land as artifacts, timeline events and findings in the case database, not as loose CSV/TXT files

  • Specialized images, not a monolith — heavy engines live in dedicated worker images (worker-plaso, worker-malware, worker-memory, …) pulled from GHCR

  • Pinned, not latest — every tool and rule set is version-pinned and hash-verified, and its version is recorded in each ToolRun

  • Verified rule provenance — every YARA / Sigma source (SigmaHQ, YARA Forge and community sets) is pinned by tag or commit and hash-verified file by file; no rule is ever silently overwritten

  • Offline & isolated — workers run without network, with dropped capabilities and resource limits

✅ v0.1 — Core + Evidence Manager

  • Case & Evidence models

  • CLI: case create, cases list, case get, evidence add/list/get/verify

  • MCP: create_case, list_cases, get_case, register_evidence, list_evidence, get_evidence, verify_evidence

  • SQLite database with provenance

  • SHA-256 hashing + integrity verification on evidence

  • Structured logging with correlation IDs

  • Docker + CI/CD

✅ v0.1.5 — Host Wrapper + Container Orchestration

  • Docker SDK integration — orchestrate forensic containers from the host

  • One container per investigation, evidence mounted read-only

  • CLI: container create/list/get/start/stop/exec/logs/remove

  • MCP: create_container, list_containers, get_container_info, start_container, stop_container, exec_in_container, container_logs, remove_container

  • Persistent workspaces at ~/.defair/workspaces/

✅ v0.2 — Windows foundation + MCP analysis

  • Dissect integration (host discovery, artifact identification)

  • 13 EZ Tools with BaseTool wrappers + normalizers: MFTECmd, EvtxECmd, RECmd, PECmd, AmcacheParser, AppCompatCacheParser, JLECmd, LECmd, RBCmd, SBECmd, SrumECmd, WxTCmd, SQLECmd

  • Normalization layer (BaseNormalizer → unified artifact schema)

  • Evidence discovery with automatic tool recommendations

  • MCP tools: discover_evidence, analyze_evtx, analyze_mft, analyze_registry, analyze_prefetch, analyze_amcache, analyze_shimcache, analyze_jumplist, analyze_lnk, analyze_recyclebin, analyze_shellbags, analyze_srum, analyze_wintimeline, analyze_sqlite

  • Tested on HackTheBox DFIR challenges (Jingle Bell, Recollection)

✅ v0.3 — Detection + Timeline + MCP hunting

  • Hayabusa v4.1 integration (4000+ Sigma rules, MITRE ATT&CK mapping)

  • Timeline Engine — unified timeline over all artifacts (summary, search, export CSV/JSONL)

  • Findings Engine — auto-created from Hayabusa high/critical detections (FND-NNN)

  • IOC search across all artifacts (description, data, hostname, username)

  • Hunting orchestration (hunt_evtx → detect → normalize → findings)

  • CLI: defair hunt, defair timeline, defair findings, defair search

  • MCP: hunt_evtx, build_timeline, search_timeline, list_findings, search_ioc

✅ v0.3.1 — Prefetch analysis fix

  • Replaced PECmd (Windows-only) with a cross-platform Prefetch parser based on libscca

  • analyze_prefetch MCP tool works end-to-end in Linux containers

✅ v0.3.5 — Mass YARA + Sigma scanning

  • YARA mass scanner on mounted evidence (files, memory dumps, disk images)

  • Sigma mass scanner via Hayabusa on all EVTX sources

  • Default rule sets embedded in the container (YARA community rules + Hayabusa Sigma rules)

  • Custom rules mounting: bind-mount /rules/yara/ and /rules/sigma/ for custom rules

  • Scan results normalized as Findings with severity, confidence, and MITRE mapping

  • MCP: scan_yara, scan_sigma — CLI: defair scan yara, defair scan sigma

  • 212 tests, 16 tool wrappers

✅ v0.3.6 — Hardening + reproducible images

Prerequisite: before adding more engines, make the platform match its own principles.

  • No shell via MCP, for real: exec_in_container removed from MCP (or gated behind an explicit mcp.allow_exec: false config flag, off by default) — kept in the CLI for humans

  • run_tool (MCP + CLI) — run a registered tool with validated arguments, recorded as a ToolRun (replaces arbitrary exec for agents)

  • Evidence allowlist — create_container only mounts paths under configured evidence roots; image restricted to ghcr.io/joblinours/defair*

  • Container hardening — cap_drop: ALL, no-new-privileges, network_mode: none, CPU / memory / PIDs limits, read-only rootfs + tmpfs

  • Dockerfile: multi-stage build (downloads in a builder stage, no compilers in the runtime image)

  • Pinned supply chain: EZ Tools and Hayabusa pinned by version + SHA-256; tool versions stored in every ToolRun (rule sets: see v0.3.7)

  • Analyst shell (CLI only): defair container shell <case> — interactive session in the case container, evidence :ro, never exposed via MCP (≈ hecatrace shell)

✅ v0.3.7 — Raijin scan engine (vendored) + verified rule sets

Raijin (Rust, YARA-X + sigma-rust) is integrated in-tree (engines/raijin/) and is the single engine behind scan_yara, scan_sigma and scan_evidence.

Engine

  • Raijin source vendored in engines/raijin/ as a DEFAIR-maintained fork (changes listed in engines/raijin/LICENSING.md), built from source with a pinned Rust toolchain in the image's raijin-build stage

  • Cold scan of the target folder only: --lab --no-procs --scan-all-files — never live processes, never other host drives

  • Replaces yara-python + the unpinned Yara-Rules/rules bundle for YARA, and Hayabusa for mass Sigma scanning

  • Native layout detection (KAPE, Velociraptor, plain mount), original Windows path and event time kept, one detection per matched event

  • Raijin JSONL now carries a structured rule reference (engine, name, id, namespace, file, tags, level) → normalizer → artifacts + findings (Sigma level / YARA score → severity, MITRE techniques from Sigma tags)

Rule sources — 13 pinned sources, ~35,000 loadable rules

Engine

Sources

Profile

YARA

YARA Forge core

precise

YARA

YARA Forge full, Elastic protections-artifacts, ESET malware-ioc, ReversingLabs, Malpedia signator-rules, Neo23x0 signature-base, Trellix ATR

broad

Sigma

SigmaHQ core + emerging-threats add-on

precise

Sigma

SigmaHQ all rules, mdecrevoisier SIGMA-detection-rules, LOLRMM

broad

Integrity — every source pinned and verified

  • Sources declared in src/defair/rules/sources.yaml, pinned in src/defair/rules/lock/rules.lock: release tag or commit SHA (never a branch), archive SHA-256, license, and a per-file SHA-256 manifest (lock/manifests/<source>.json) — shipped inside the Python package

  • defair rules lock --refresh re-pins (release archives checked against GitHub's published digest); bumping the lock is a reviewed commit

  • defair rules sync (image build) installs each source into /opt/defair/rules/<engine>/<source>/ and verifies every file — any missing, extra or modified file aborts the build

  • Re-verified before every scan: a store that differs from the lock → scan refused

  • raijin-util validate runs per source at build (VALIDATION.txt); the build fails if a source has no loadable rule

  • raijin-util update / upgrade disabled — no unpinned latest downloads

Collision handling — no rule silently overwritten

  • No basename flattening: each source keeps its upstream tree, so same-named files in different folders or sources all survive

  • Per-run signature tree assembled as NN_<source> symlinks in lock order: first source wins on a duplicate Sigma id, each YARA source loads in its own namespace

  • CONFLICTS.json per profile: identical Sigma duplicates, conflicting Sigma rules (same id, different content — first source wins) and YARA rule names shipped by several sources (defair rules conflicts, MCP list_rule_conflicts)

  • YARA hits of the same rule from several sources are merged into one artifact listing every source

Provenance in results

  • Every finding's detection_refs: engine, source, repo, pinned ref, rule id / name, rule file path + SHA-256, license

  • Custom rules (/rules/yara/, /rules/sigma/ or --yara-rules-dir / --sigma-rules-dir) linked as 99_custom, hashed into the ToolRun, tagged provenance: custom

  • /opt/defair/rules/NOTICE lists every embedded rule set with its license (DRL 1.1, Elastic License 2.0, CC BY-SA 4.0, MIT, BSD…)

  • CLI: defair scan yara|sigma|evidence --profile precise|broad, defair rules status|verify|conflicts|lock|sync

  • MCP: scan_yara, scan_sigma, scan_evidence, get_ruleset_info, list_rule_conflicts

  • Fixed: scan findings were never created (the service read artifact_count instead of artifacts_produced)

  • Hayabusa kept for hunt_evtx timeline enrichment, pinned by version + SHA-256

  • Deferred: offline rule updates from a verified bundle (defair rules update --bundle) — rules are updated by re-pinning the lock and rebuilding the image

✅ v0.3.8 — Preprocessing & normalization pipeline

Inspired by ArtefactProcessor / PyTriage, adapted to DEFAIR's provenance model.

  • Two-stage normalization: tool output → normalized JSONL (/workspace/normalized/<artifact_type>/<RUN>.jsonl, SHA-256 recorded in normalized_files) → batched bulk insert (2,000 rows per batch, numbering computed once per run instead of a COUNT + commit per row)

  • Deterministic artifact ids (uuid5(run_id, record_key)) — re-normalizing the same output yields the same ids and ART-NNN numbers, so findings stay linked

  • defair normalize replay --case rebuilds a case from its JSONL (files whose hash changed are refused); defair normalize rerun RUN-xxx re-normalizes from the raw tool output after a normalizer fix; defair normalize stats RUN-xxx

  • Common envelope on every artifact (provenance): tool + pinned version, run, evidence, source file, host path, channel, record id, raw value of any unparseable timestamp

  • Timestamps: ISO 8601 UTC with full source precision (EZ Tools' 7 fractional digits kept); an unparseable time stays null with a reason — never replaced by "now"; ambiguous dd/mm vs mm/dd formats are refused

  • Timeline fields on every event: timestamp_desc (Created ($SI), Last Executed, Event Logged…) and message; defair timeline export --format timesketch (MCP export_timeline)

  • Generic EVTX flattening: System + EventData / UserData → flat keys, for EvtxECmd's Payload and native records

  • EventID knowledge base (src/defair/data/evtx_catalog.yaml, 89 events: Security, System, Sysmon, PowerShell, RDP, Task Scheduler, Defender, WMI, BITS, USB): channel-aware type, category, description and MITRE techniques — add an event without code

  • Pure-Python fallbacks: evtx_native (pyevtx-rs) for EvtxECmd, lnk_native (LnkParse3) for LECmd — run automatically when the tool fails, as their own ToolRun linked by fallback_of

  • Per-run counters in tool_runs.normalization_stats: rows read, normalized, skipped, errors, unparseable timestamps, reasons

  • Schema migrations (PRAGMA user_version): databases from earlier versions upgrade in place

  • MCP: normalize_replay, normalize_rerun, get_normalization_stats, export_timeline

  • Not copied from ArtefactProcessor: lossy dd/mm/YYYY HH:MM:SS timestamps, datetime.now() substituted for missing times, silently swallowed exceptions

  • Deferred: JumpList pure-Python fallback

✅ v0.4 — Evidence sources + Orchestration + MCP profiles

🎯 Milestone: MVP MCP — one call runs a full Windows investigation, whatever the evidence format.

Evidence sources (src/defair/sources/)

  • Detection: KAPE (folder / VHDX), Velociraptor, FastIR, UAC, mounted filesystems, log folders, disk images (E01 / Ex01 / VMDK / VHD / VHDX / QCOW2 / raw), ZIP (plain, ZipCrypto, AES), Generaptor, DFIR-ORC — plus artifacts identified by content (magic bytes) when a collector renamed them

  • Collection folders registered as evidence with a tree hash (verify detects any added / removed / modified file)

  • Preparation (defair evidence prepare, MCP prepare_evidence): collections used in place, read-only; ZIP extracted (zip-slip and archive-bomb guards, ZIP-in-ZIP); Generaptor decrypted (RSA-OAEP + AES); DFIR-ORC decrypted with ANSSI's orc-decrypt (vendored in engines/orc-decrypt/, LGPL-2.1) + nested 7z; disk images carved with Dissect — no mount, no privilege, no partition offset to compute — plus a host profile artifact

  • manifest.json records every derived file (origin, size, SHA-256); secrets (archive password, key passphrase) travel through the exec environment, never on a command line, never stored; private keys mounted read-only at /keys (container create --keys, validated against container.key_roots)

Orchestration (src/defair/orchestrator/, src/defair/profiles/)

  • Declarative profiles: windows-triage, windows-full, ransomware, persistence, registry-only, scan-only

  • DAG executor: dependencies, bounded parallelism (orchestrator.max_parallel), per-step timeout, retry with backoff, optional steps, cancellation (running tools are killed), resume (completed steps kept)

  • Engines: auto = EZ Tools → pure-Python fallbacks → Dissect plugins on the original image / collection; ez = no Dissect; dissect = Dissect plugins only (dissect_plugin tool + generic normalizer, same artifact types as EZ Tools)

  • Background profile runs PRUN-NNN: detached worker, status / list / cancel / resume, dead-worker detection

  • run.json written after every step and always at the end — even on failure: DEFAIR version, profile, engine, evidence preparation, each step's tools, pinned versions, fallbacks used, durations, errors, rule lock

  • Concurrency fixes: RUN / ART / FND numbers reserved under a lock (parallel steps collided on RUN-NNN)

  • CLI: defair profile list|show, defair run start --case CASE-xxx --evidence EVD-001 --profile windows-triage (≈ hecatrace run -v -e), defair run status|list|cancel|resume

  • MCP: prepare_evidence, list_profiles, get_profile, run_profile, analyze_evidence, get_run_status, list_runs, cancel_run, resume_run

  • Not yet: Linux disk images (UAC collections are scanned with Raijin only), BitLocker-encrypted volumes

✅ v0.4.1 — Investigation ergonomics (from the first real E01 run)

  • --case accepts the case name everywhere (case-insensitive; ambiguous names refused), as well as CASE-YYYY-NNN or the id

  • findings get explains every match: evidence file (real path in the container), event (EventID / RecordID / Computer) or YARA offset, and the exact field values / pattern that hit the rule; rule id, rule file and defair rules show <source> <path>; --all, --json (with the raw EVTX event / hex context)

  • artifacts get ART-NNN: every field and the full data, provenance (tool, pinned version, RUN-NNN, input), linked findings, normalized JSONL, the complete EVTX event read back from the log by record id, a hex dump around each YARA match, --context MIN for the surrounding timeline

  • artifacts list: --contains (any field), --tool, --host, --user, --severity, --since/--until, --run, --asc, paging (--offset), --json

  • Container logs: every command (arguments with secrets redacted, output, exit code, duration), tool run (command, exit code, stderr), profile step and finding is logged as JSON to /workspace/logs/defair.log; the container's PID 1 follows it, so docker logs <container> shows everything. Console logs moved to stderr (stdout carries results only)

  • container create refuses a workspace the host user cannot write (created by DEFAIR < 0.3.6 running as root) with the chown fix

  • Hayabusa 4.x: dfir-timeline syntax, run from /opt/hayabusa, abbreviated levels (crit, med) mapped — critical detections now become findings

  • MCP: get_artifact, get_finding, show_rule; list_artifacts gains every filter + paging

✅ v0.4.5 — Windows coverage completion

Fixes first

  • MFTECmd $J output is normalized as windows.usn.journal_entry (it was a file entry without timestamp); -m $MFT resolves parent paths; Dissect usnjrnl records use the same type

  • SrumECmd: one artifact type per SRUM table (app resource usage, network usage / connectivity, energy, push notifications) instead of network_usage for everything

  • MCP analyze_srum forwards registry_hive; MFTECmd --bdl takes the drive letter only

Remaining EZ Tools (pinned by SHA-256 like the others)

  • RecentFileCacheParser (Windows 7 program execution), SumECmd (User Access Logging — who reached which Windows Server role, from where), bstrings (pattern search in any binary; runs with a pseudo-terminal on stdin), rla (dirty hive replay)

  • Profile steps can read another step's output: input: step:hives_replay (+ input_fallback) — rla replays the .LOG1/.LOG2 of dirty hives into the workspace, RECmd parses the clean copies

NTFS depth — native parsers on dissect.ntfs, no mount, E01 / VMDK / VHDX / raw

  • $I30 INDX slack (indx_native, INDXRipper approach): every directory of every NTFS volume; $FILE_NAME remnants of deleted / renamed files with their four $FN times; entries still live (allocation or $INDEX_ROOT) are dropped

  • $LogFile (logfile_native): file names linked / unlinked, FILE records created / freed, $FILE_NAME created / removed; every other operation is counted as not decoded in the run statistics, never guessed

  • USN journal in windows-triage

Extra Windows artefacts (pure-Python parsers, Dissect plugins as fallbacks)

  • Defender MPLog (detections, processes Defender measured, SDN file hashes, exclusions), PowerShell ConsoleHost_history, Scheduled Task XML (actions, triggers, principal — defusedxml), WebCache (IE / legacy Edge history incl. Explorer file:// accesses, downloads, cookies), RDP bitmap cache (tiles rebuilt as PNG + a collage), IIS W3C logs

  • Local times without a zone (task registration, some MPLog lines) are kept as text, never assumed UTC

Typed EVTX views — src/defair/data/evtx_views.yaml, on top of the EventID catalog

  • logons, process_creation, services, scheduled_tasks, rdp, powershell, account_changes, log_cleared, defender, network_shares, kerberos_ntlm — filtered, column-projected views over the EVTX artifacts of any parser (EvtxECmd, native, Dissect); schema v4 indexes the EventID

  • CLI defair evtx views / defair evtx view logons --case X [--since/--until/--host/--user/--event-id]

Chainsaw — second Sigma engine, pinned (2.16.5, SHA-256 = GitHub digest), on the verified DEFAIR rule store (never its own bundle): rule file, SHA-256, source, pinned ref resolved by Sigma id; findings like Raijin's — defair hunt --engine chainsaw --rule-profile precise|broad

Host profile — hostname, domain, OS build, architecture, timezone, install date, users, IPs, installed applications (Dissect, also on collections), registry time zone / network profiles / USB devices / services, computer names in the logs — every fact with its source, disagreements listed as conflicts, stored as one windows.system.host_profile artifact per evidence (≈ Hecatrace systeminfo.txt) — defair host profile, profile action host_profile

Strings + watchlists

  • Strings extraction (strings_native): ASCII + UTF-16LE of pagefile.sys / swapfile.sys, streamed through Dissect (never copied), written as a TSV index; hiberfil.sys reported as skipped (compressed — v0.8), unallocated space in v0.5

  • Watchlists: built-in (offensive_tools, lolbins, rmm, exfiltration) + per case (/workspace/watchlists/*.yaml), literal or regex terms; one batch search over the normalized artifacts (hits → ART-NNN), the strings index and the raw collection files (ASCII + UTF-16), ripgrep-backed (pinned) with a pure-Python fallback; JSON report + optional findings — defair watchlist list|show|search

  • MCP: analyze_usn, analyze_ual, hunt_chainsaw, list_evtx_views, get_evtx_view, get_host_profile, list_watchlists, search_watchlist

  • Not yet: hiberfil decompression (v0.8), JumpList pure-Python fallback

✅ v0.5 — Supertimeline

Worker images — heavy engines out of the main image

  • workers: in defair.yaml (image pinned by version tag — latest is refused — plus memory / CPU / PIDs / timeout)

  • The case container has no Docker socket: the host starts a short-lived job container next to it (services/worker_service.py) with the case container's evidence (read-only, same paths) and workspace mounts — never /keys — and the same hardening, network always none

  • The job is /workspace/jobs/<RUN>/job.json, written by the case container: argv lists only, never a shell; python -m defair.workers.entry runs it, writes status.json after every step, logs to the case log (docker logs shows the job); SIGTERM kills the running tool and records cancelled

  • CI builds a matrix of images: ghcr.io/joblinours/defair and ghcr.io/joblinours/defair-worker-plaso, each checked after publication (python -m defair.workers.check)

worker-plaso image (docker/worker-plaso.Dockerfile)

  • Plaso 20260720 and all 73 Python dependencies installed with pip --require-hashes from docker/worker-plaso.requirements.txt

  • The Sleuth Kit 4.15.0 + libewf-legacy 20140816 compiled from tarballs pinned in docker/worker-plaso.checksums.sha256 (E01 / Ex01 read directly); no compiler in the runtime image

Jobs — defair supertimeline start --case X --evidence EVD-001 --mode plaso|bodyfile|both|unallocated|all

  • plaso: log2timeline into /workspace/plaso/<EVD>.plaso, then psort -o json_line. The storage is reused (only psort runs again) when it was built from the same evidence hash with the same parsers and time zone and still matches its recorded SHA-256; built differently → refused unless --overwrite; a half-built storage is deleted

  • bodyfile (disk images): mmls, then fls -r -m per partition — the filesystem type is tried explicitly (-f ntfs, fat, ext…, the mmls description first) instead of TSK's autodetection; the bodyfile is parsed directly (one event per distinct MACB time, like mactime, epoch = UTC)

  • unallocated: blkls per partition streamed into the strings extractor → /workspace/strings/<EVD>/unallocated-<n>.tsv, searchable with search_watchlist

Timeline Engine

  • Schema v5: timeline_events table (no ART-NNN), composite (case_id, timestamp) index, FTS5 index on message + file name

  • Streamed import (bounded memory — 554,138 events of a real Windows 10 E01 imported in 60 s with 173 MiB), normalized JSONL + SHA-256, deterministic ids (re-import replaces), normalize replay restores events too, normalize rerun RUN re-imports a supertimeline run

  • Timestamps at source precision (Plaso date_time: FILETIME keeps 100 ns); Plaso's "not a time" stays null with the raw value

  • timeline summary|search|export merge artifacts and events in time order: --sources artifacts,plaso,tsk, --parser, --type, --offset; text search = LIKE on artifacts, FTS5 phrase on events; exports are streamed with no row limit (the former 100,000 cap is gone)

  • run manifest per job (/workspace/jobs/<RUN>/manifest.json): spec, runner steps, worker image + digest, Plaso / TSK versions, storage reuse and hash, import counts

  • MCP: build_supertimeline, get_supertimeline_status (imports when the job ends), cancel_supertimeline; search_timeline / export_timeline extended to Plaso and Sleuth Kit sources

  • Not a profile step: profile runs execute inside the case container, which cannot start containers — the supertimeline is started from the host (CLI / MCP)

📋 v0.6 — Reporting + REST API

  • Forensic reports (Markdown, HTML, JSON) — findings, host profile, timeline highlights, tool runs and versions

  • Case export — per-category CSV/JSONL tree for human review (Timeline Explorer, spreadsheets)

  • Optional export connectors: Timesketch (timeline) and OpenSearch (bulk JSONL) — the case database stays the source of truth

  • REST API (FastAPI + OpenAPI)

  • MCP: generate_report, export_case

📋 v0.7 — Malware & document triage

  • Dedicated worker-malware image (offline, no network)

  • File extraction from evidence into the workspace with hash + source provenance (extract_file)

  • capa (capabilities), FLOSS (obfuscated strings), pefile (PE metadata), ssdeep / TLSH (fuzzy hashing)

  • Documents: oletools, oledump, pdfid, pdf-parser, ExifTool metadata

  • ClamAV with a pinned, offline signature database

  • YARA triage of extracted files through Raijin (same pinned rule sets as v0.3.7)

  • Results normalized as artifacts + findings (MITRE ATT&CK from capa)

  • MCP: extract_file, triage_file, scan_clamav

📋 v0.8 — Memory forensics

  • Volatility 3 in a dedicated worker-memory image (pinned symbol tables for offline use)

  • Profiles: memory-triage (pslist/pstree, cmdline, netscan, malfind, svcscan, dlllist)

  • Memory artifacts correlated with disk artifacts in the timeline and findings

  • YARA + strings on memory dumps

  • MCP: analyze_memory, run_profile memory-triage

📋 v0.9 — Carving + deep disk

  • Dedicated worker-carving image: bulk_extractor (emails, URLs, IPs, credit cards…), PhotoRec / foremost / scalpel, binwalk

  • Carved items registered as derived evidence (parent evidence + offset + hash)

  • Image verification: ewfverify, hashdeep integrated into verify_evidence

  • MCP: carve_evidence, run_bulk_extractor

  • 🎯 Milestone: coverage parity with all-in-one DFIR toolboxes — with structured results, provenance and MCP access

📋 v0.10+ — Beyond parity

  • Email: PST/OST/MBOX parsing, attachment extraction, OCR on attachments

  • Cloud & AD sources (ArtefactProcessor plugins): DFIR-O365RC, Google Workspace, ADTimeline, ADAudit

  • Zeek / TShark / Suricata — network DFIR

  • Linux DFIR — journald, SSH, cron, systemd, Docker artifacts; configurable log globs (auth, audit, nginx, apache…); Sigma on Linux logs via Raijin

  • Web UI

  • RBAC, audit trail, SBOM + image signing

📋 v1.0 — DEFAIR

  • Complete forensic workflow

  • Windows + Linux + Memory + Network

  • Stable MCP + API

  • Reproducible reports

  • Offline installation

  • Production-ready security

See defair_Roadmap.md for the full detailed roadmap.


Forensic engines

Engine

Purpose

Worker image

Phase

Status

Dissect

Image carving without mount (v0.4), plugins as parsing fallback / engine=dissect

defair

v0.2 / v0.4

✅

orc-decrypt (ANSSI)

DFIR-ORC archive decryption

defair

v0.4

✅

EZ Tools (13)

Windows artifacts (MFT, EVTX, Registry, Amcache, LNK, SRUM, ...)

defair

v0.2

✅

Hayabusa

EVTX hunting timeline (hayabusa-rules, distinct provenance)

defair

v0.3

✅

YARA (yara-python)

File pattern matching — replaced by Raijin in v0.3.7

—

v0.3.5

⛔

Raijin (vendored)

YARA-X + Sigma cold scanner, 13 pinned rule sources

defair

v0.3.7

✅

EZ Tools (remaining)

RecentFileCacheParser, SumECmd, bstrings, rla

defair

v0.4.5

✅

NTFS native (dissect.ntfs)

$I30 slack, $LogFile, USN Journal

defair

v0.4.5

✅

Chainsaw

Second Sigma engine on the pinned rule store

defair

v0.4.5

✅

ripgrep

Watchlist / IOC batch search

defair

v0.4.5

✅

Plaso

Supertimeline, multi-source timestamp normalization

worker-plaso

v0.5

✅

Sleuth Kit

Bodyfile filesystem timeline, unallocated space (blkls)

worker-plaso

v0.5

✅

capa / FLOSS / pefile

Malware capabilities, obfuscated strings, PE metadata

worker-malware

v0.7

📋

oletools / Didier Stevens suite / ExifTool

Office, PDF and metadata triage

worker-malware

v0.7

📋

ClamAV

AV scanning (pinned offline signatures)

worker-malware

v0.7

📋

Volatility 3

Memory forensics (processes, network, DLLs, persistence)

worker-memory

v0.8

📋

bulk_extractor / PhotoRec / foremost / scalpel

Feature extraction and file carving

worker-carving

v0.9

📋

Zeek / TShark / Suricata

Network traffic analysis, IDS

worker-network

v0.10+

📋

Out of scope: acquisition tools (dc3dd, imaging) — DEFAIR analyses evidence, it does not acquire it.


Development

# Run tests
pytest tests/ -v

# Lint
ruff check src/ tests/

# Run CLI
defair --help

# Run MCP server
defair-mcp

Adding a new MCP tool

  1. Add the service function in src/defair/services/

  2. Add the CLI command in src/defair/cli/

  3. Add the MCP tool in src/defair/mcp_server/server.py with @mcp.tool()

  4. Add tests (unit + CLI + MCP integration)

  5. Both CLI and MCP must call the same service function


License

MIT


DEFAIR: a reproducible DFIR platform capable of orchestrating forensic engines, unifying their results, and exposing the investigation to a human or an agent via CLI, API, and MCP.

Available Tools

7 tools
create_caseCreate CaseC

Create a new forensic investigation case.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName of the case (e.g. "Incident host-01 ransomware").
descriptionNoOptional longer description of the case.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it says nothing beyond the act of creation: no persistence guarantees, no required auth/permission level, no indication of whether the case is immediately visible to get_case/list_cases, and no side effects. For a mutation tool this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the verb-resource pair front-loaded and zero filler. It is efficient, though its brevity is partly under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the schema fully covers the two simple parameters. However, as an unannotated mutation tool it omits permissions, side effects, and any post-creation workflow hints, leaving the definition minimally but not fully adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents both 'name' and 'description' including an example and a default. The description adds no parameter meaning beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a ... case') and the domain ('forensic investigation'), which is enough to distinguish it from the read-oriented siblings (list_cases, get_case, verify_evidence). It does not explicitly name or contrast the nearest alternative, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to create a case versus reusing an existing one, no prerequisites (e.g., authentication or permissions), and no mention of sibling tools such as register_evidence that would follow. The agent must infer context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_caseGet CaseB

Get details of a specific forensic case.

ParametersJSON Schema
NameRequiredDescriptionDefault
case_idYesThe case number (e.g. "CASE-2026-001") or internal UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Get' weakly implies a read, but the description never states that this is a non-mutating lookup, what permissions are required, or whether case data is sensitive/auditable. With zero annotation coverage this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no waste, front-loaded with the verb and resource. Appropriate size for a one-parameter lookup tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be described, and the single parameter is fully documented in the schema. But with no annotations and no usage guidance, an agent still lacks context on safety profile and when to prefer this over list_cases/get_evidence. Minimum viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the case_id description already provides format examples (CASE-2026-001, internal UUID). The description adds nothing about the parameter, so the baseline of 3 applies – the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('details of a specific forensic case'), clear enough to distinguish from siblings like list_cases (plural, listing) and create_case (mutation). However it doesn't explicitly contrast with get_evidence or explain what 'details' includes, leaving some ambiguity about sibling selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no mention of alternatives. The singular 'a specific forensic case' weakly implies you need a known case_id, but nothing tells the agent when to reach for this versus list_cases or get_evidence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_evidenceGet EvidenceC

Get details of a specific evidence item.

ParametersJSON Schema
NameRequiredDescriptionDefault
evidence_idYesEvidence number (e.g. "EVD-001") or internal UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden but only implies a read via 'Get'. It says nothing about whether the evidence must exist, error behavior for unknown IDs, any access restrictions, or what happens with the two accepted ID formats. Minimal disclosure for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words and the verb/resource established immediately. It is efficient, though almost too terse to carry any extra value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and there is only one well-documented parameter. The definition is adequate for a simple lookup but omits error/edge behavior and sibling differentiation that would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter is documented with concrete examples ("EVD-001" or internal UUID). The description adds no additional meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('a specific evidence item'), which is clear and unambiguous. However, it does nothing to distinguish itself from siblings like get_case or list_evidence/verify_evidence, leaving the agent to infer from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus list_evidence, verify_evidence, or register_evidence. The description offers no context, prerequisites, or exclusions, so an agent must guess that this is the single-item lookup counterpart to list_evidence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_casesList CasesB

List all forensic cases.

Returns a list of cases with their case number, name, status, and creation date. Cases are ordered by creation date (newest first).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add real behavioral context: the ordering (newest first) and the fields returned. However, it is silent on whether the full result set is returned unbounded, whether pagination or limits exist, and whether any permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose. The second sentence largely restates the output schema's fields, which is mild redundancy but not enough to hurt readability or comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained in prose, and the empty input schema means there are no parameter gaps to cover. The only missing piece is whether results are bounded or paginated, which is a minor gap for a zero-parameter list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate beyond confirming the operation takes no input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List all forensic cases') and the scope ('all') implicitly distinguishes it from get_case, which retrieves a single case. It is unambiguous but never explicitly contrasts itself with the sibling read tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus get_case, and no mention of whether the result set is filtered, paginated, or capped. The agent must infer that 'all' means no filtering and no pagination control exists given the empty parameter schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_evidenceList EvidenceB

List registered evidence items.

ParametersJSON Schema
NameRequiredDescriptionDefault
case_idNoOptional — filter by case number (e.g. "CASE-2026-001") or UUID. If omitted, returns all evidence across all cases.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It doesn't state whether this is a read-only operation, whether it paginates, what the return volume might be, or any auth requirements. For a list tool with no annotations, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is concise and front-loaded, but it's so terse that it sacrifices useful context. It earns points for zero waste, but doesn't provide enough substance to be considered well-structured for a list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained. The single optional parameter is fully documented in the schema. However, with no annotations and no behavioral context, the description is barely adequate for an agent to understand when and how to use this tool versus siblings like get_evidence or list_cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — the case_id parameter is fully documented in the schema including its optional nature, accepted formats (case number or UUID), and default behavior. The description adds no parameter information beyond what's already in the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource: listing registered evidence items. This distinguishes it from get_evidence (singular retrieval) and register_evidence (creation). However, it doesn't explicitly contrast with list_cases, the closest sibling, leaving some sibling differentiation implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (you call it to see a list of evidence), but provides no explicit when-to-use guidance, no exclusions, and doesn't name alternatives like get_evidence for single-item retrieval. The optional case_id parameter's behavior is documented in the schema, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_evidenceRegister EvidenceA

Register a new piece of evidence in a forensic case.

Computes SHA-256 hash of the file and stores metadata. The original file is never modified or copied.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to the evidence file.
case_idYesCase number (e.g. "CASE-2026-001") or UUID.
evidence_typeNoType of evidence — one of: disk_image, memory_dump, logs, triage_archive, pcap, other.other

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and discloses important side effects: it computes a SHA-256 hash, stores metadata, and explicitly states the original file is never modified or copied. It still omits permissions, duplicate-registration behavior, and error conditions, but the core mutation and file-safety behavior are clearly conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with the purpose front-loaded, followed by concise behavioral details. Every sentence adds useful information without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, complete parameter descriptions, and an existing output schema, the description covers the essential action and key side effects. It is nearly complete, though it could mention prerequisites such as whether the case must already exist or what happens on duplicate registration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters, including case_id format, path, and evidence_type enum values. The description adds no parameter-level meaning beyond what the schema provides, which matches the baseline of 3 for fully covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Register') and resource ('a new piece of evidence in a forensic case'), making the action immediately identifiable. It distinguishes itself from sibling tools like get_evidence, list_evidence, and verify_evidence by focusing on registration of new evidence rather than retrieval or verification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use guidance, prerequisites, or alternative tools. It implies registration is for adding new evidence, but does not say when to prefer it over verify_evidence or how to handle existing evidence, leaving usage conditions entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_evidenceVerify EvidenceA

Verify evidence integrity by re-computing its SHA-256 hash.

Compares the current file hash against the hash stored at registration. This is a critical forensic operation to detect evidence tampering.

ParametersJSON Schema
NameRequiredDescriptionDefault
evidence_idYesEvidence number (e.g. "EVD-001") or internal UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the core behavior (recompute hash, compare to the value stored at registration) and implies a non-destructive read, but never states permissions, whether the result is persisted or logged, or what happens on a mismatch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the mechanism front-loaded and no filler. The phrase "critical forensic operation" is mildly editorial but harmless, so structure is strong though not perfect.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return shape need not be explained, and the one parameter is fully covered. The description is sufficient to call the tool, though it leaves the mismatch behavior and read/write semantics implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single evidence_id parameter is already fully documented in the schema (evidence number or UUID). The prose adds no parameter-level detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource (verify evidence) and even names the mechanism (re-computing its SHA-256 hash) plus the goal (detect tampering). This clearly distinguishes it from the read-only siblings get_evidence and list_evidence, which retrieve rather than validate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies usage ("critical forensic operation to detect evidence tampering"), which tells an agent the context in which it matters. However, it names no alternatives or preconditions, e.g. whether get_evidence must be called first or when verification is unnecessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.0
    • First observedcreate_case
    • First observedget_case
    • First observedget_evidence
    • First observedlist_cases
    • First observedlist_evidence
    • First observedregister_evidence
    • First observedverify_evidence

TDQS

A3.6/5.0

Scored across 7 tools

Disambiguation5/5

Each tool targets a distinct resource (case vs. evidence) and action (list, create, get, register, verify). The purposes are clear and non-overlapping, with no ambiguity in selection.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern: list_cases, create_case, get_case, register_evidence, list_evidence, get_evidence, verify_evidence. This is predictable and easy to parse.

Tool Count5/5

Seven tools are well-scoped for a forensic case management server, covering essential operations without redundancy. Each tool earns its place in the set.

Completeness4/5

The surface covers core case and evidence lifecycle operations, including creation, retrieval, listing, and integrity verification. However, it lacks update and delete operations for cases and evidence, which are common CRUD needs that could be required for full lifecycle management.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers