DEFAIR MCP Server
This MCP server provides forensic case and evidence management for DEFAIR investigations.
Manage cases: Create, list, and retrieve forensic investigation cases.
Register evidence: Add evidence files to a case with automatic SHA-256 hashing and metadata (type, path); originals are never modified.
List evidence: View all registered evidence, optionally filtered by case.
Inspect evidence: Get details of a specific evidence item by number or UUID.
Verify integrity: Re-hash evidence and compare against the stored hash to detect tampering.
Allows orchestration of forensic Docker containers, including creating containers linked to investigation cases, listing, starting, stopping, removing them, executing commands inside containers, and retrieving container logs.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@DEFAIR MCP Serverregister /evidence/host01.E01 as disk image evidence for CASE-2026-001"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DEFAIR
Digital Forensics & Incident Response platform — MCP-first, containerized, modular.
What is DEFAIR?
DEFAIR is a reproducible DFIR platform that orchestrates multiple forensic engines, unifies their results, and exposes the investigation to a human or an AI agent via CLI, API, and MCP (Model Context Protocol).
It is not "a giant Docker container with 50 forensic binaries". It is a forensic orchestration platform where external tools are specialized engines managed through a common architecture.
Analyst / AI Agent
│
┌───────────┴───────────┐
│ │
CLI MCP
│ │
└───────────┬───────────┘
│
Service Layer
│
┌───────┴────────┐
│ Orchestrator │
└───────┬────────┘
│
┌─────────────────┼─────────────────┐
│ │ │
Evidence Manager Tool Registry Job Engine
│ │ │
└─────────────────┼─────────────────┘
│
Normalization Layer
│
┌──────────────┼───────────────┐
│ │ │
Timeline Findings IOC
│ │ │
└──────────────┼───────────────┘
│
Reports / ExportCore principles
MCP-first — every capability is exposed via CLI and MCP simultaneously
Read-only on evidence — source files are never modified
Hash & provenance — every result is traceable back to its source
Reproducible — every execution is logged, versioned, and replayable
No shell via MCP — the MCP exposes forensic operations, not arbitrary commands
Offline-first — designed to work without internet access
Related MCP server: Forensic Artifact Investigator MCP Server
Quick start
Installation
# Clone
git clone https://github.com/joblinours/defair.git
cd defair
# Create venv and install (dev includes pytest, ruff, etc.)
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
# Verify installation
defair --version
pytest tests/ -vCLI usage
# Create a forensic case
defair case create "Incident 2026-09 — Host compromise"
# List cases
defair cases list
# Register evidence (computes SHA-256 automatically)
defair evidence add CASE-2026-001 /evidence/host01.E01 --type disk_image
# List evidence (all or filtered by case)
defair evidence list
defair evidence list --case CASE-2026-001
# Get details
defair case get CASE-2026-001
defair evidence get EVD-001
# Verify evidence integrity (re-hash and compare)
defair evidence verify EVD-001
# Whole investigation in one command (v0.4): collection folders, archives,
# Generaptor / DFIR-ORC, disk images — prepared, then a profile runs in the background
defair run start --case CASE-2026-001 --evidence /evidence/host01.E01 --profile auto
defair run status PRUN-001Container orchestration
DEFAIR runs forensic tools in isolated Docker containers — one per investigation.
Containers are hardened by default: all capabilities dropped, no-new-privileges, no network,
read-only root filesystem, CPU/memory/PID limits, running as your host user. Evidence can only
be mounted from the directories listed in container.evidence_roots (defair.yaml).
Hardening applies to containers created from v0.3.6 on — recreate older ones.
# Create a forensic container linked to a case
defair container create --case CASE-2026-001 --evidence /evidence/host01.E01
# List containers
defair container list
# Interactive analyst shell (CLI only — never exposed through MCP)
defair container shell defair-case-2026-001
# Execute a command inside the container
defair container exec defair-case-2026-001 ls -la /evidence
# Get container details / logs
defair container get defair-case-2026-001
defair container logs defair-case-2026-001
# Stop / remove
defair container stop defair-case-2026-001
defair container remove defair-case-2026-001MCP usage
DEFAIR exposes a MCP server over stdio — compatible with Claude Desktop, Claude Code, and any MCP client:
# Run the MCP server
defair-mcpAvailable MCP tools:
Tool | Description |
Case & Evidence | |
| Create a new investigation case |
| List all forensic cases |
| Get case details by ID or case number |
| Register evidence with SHA-256 hash |
| List evidence (optionally filtered by case) |
| Get evidence details by ID or number |
| Re-hash evidence and verify integrity |
Container | |
| Create an isolated forensic container |
| List DEFAIR containers |
| Get container details |
| Start a stopped container |
| Stop a running container |
| Remove a container |
| Run a registered tool with validated options (recorded as a ToolRun) |
| Arbitrary command — disabled unless |
| Get container logs |
Discovery & Analysis | |
| Discover forensic artifacts on mounted evidence |
| List available forensic tools |
| Check tool availability in a container |
| List past analysis runs |
| Search artifacts (text, tool, host, user, time range, run) with paging |
| Deep-inspect one artifact: data, provenance, raw EVTX event / YARA hex, timeline context |
| A finding with every match explained (file, event / offset, matched values, rule file) |
| Content of a YARA / Sigma rule from the verified store |
| Parse Windows Event Logs (EvtxECmd) |
| Parse NTFS Master File Table (MFTECmd) |
| Parse Windows Registry hives (RECmd) |
| Parse Prefetch files (PECmd) |
| Parse Amcache.hve (AmcacheParser) |
| Parse Shimcache (AppCompatCacheParser) |
| Parse Jump Lists (JLECmd) |
| Parse LNK shortcuts (LECmd) |
| Parse Recycle Bin (RBCmd) |
| Parse ShellBags (SBECmd) |
| Parse SRUM database (SrumECmd) |
| Parse Windows Timeline (WxTCmd) |
| Parse SQLite databases (SQLECmd) |
Detection & Hunting (v0.3) | |
| Run Hayabusa Sigma detection on EVTX |
| Build unified timeline summary |
| Search/filter the timeline (artifacts + Plaso / Sleuth Kit events) |
| List investigation findings |
| Search IOC across all artifacts |
Scanning & rules (v0.3.7) | |
| YARA scan of every file (Raijin, pinned rule sets) |
| Sigma scan of EVTX / Linux logs (KAPE, Velociraptor, mount) |
| YARA + Sigma in a single pass |
| Pinned rule sources, installation and integrity status |
| Duplicate / conflicting rules across sources |
Investigations (v0.4) | |
| Path → registered, prepared, profile chosen and run (background) |
| Extract / decrypt / carve an evidence (ZIP, Generaptor, DFIR-ORC, disk images) |
| Analysis profiles and their steps |
| Run a profile on an evidence in the background ( |
| Follow profile runs (steps, tools, fallbacks, errors) |
| Stop a run / resume it without re-running completed steps |
Windows coverage (v0.4.5) | |
| USN journal ($J) with MFTECmd, parent paths from the $MFT |
| User Access Logging (Windows Server) with SumECmd |
| Sigma hunt with Chainsaw on the pinned rule store |
| Typed EVTX views (logons, services, RDP, PowerShell…) |
| Host profile with a source per fact |
| Keyword / IOC watchlists over artifacts, strings, raw files |
Supertimeline (v0.5) | |
| Plaso / Sleuth Kit job in the |
| Job state; imports the events into the timeline when finished |
| Stop a supertimeline job |
Normalization & export (v0.3.8) | |
| Export the timeline (Timesketch JSONL, JSONL, CSV) |
| Rebuild a case's artifacts from its normalized JSONL |
| Re-normalize a tool run from its raw output |
| Normalization counters of a run |
Docker
# Build
docker compose build
# Run CLI
docker compose run --rm defair defair case create "Docker test"
# Run MCP server
docker compose run --rm mcpArchitecture
DEFAIR follows a triple-interface architecture: CLI, MCP, and API all call the same async Service Layer. No business logic lives in the interface layers.
CLI (click) MCP (FastMCP) API (FastAPI — planned)
│ │ │
│ run_sync() │ async direct │ async direct
└─────────┬───────────┴──────────────────────┘
│
Service Layer (async)
│
┌─────────┴──────────┐
│ │
SQLite Docker SDK
(aiosqlite) (container_service)
│
┌───────────┴───────────┐
│ DEFAIR Container 1 │ evidence :ro
│ DEFAIR Container 2 │ workspace :rw
└───────────────────────┘Project structure
src/defair/
├── config.py # YAML + Pydantic configuration
├── logging.py # Structured logging (structlog)
├── database.py # SQLite schema & connection management
├── models/ # Pydantic data models (Case, Evidence, ...)
├── services/ # Async service layer (shared by CLI & MCP)
├── cli/ # Click CLI commands
└── mcp_server/ # FastMCP server & tool definitionsTech stack
Component | Technology |
Language | Python 3.13 |
CLI | Click + Rich |
MCP server | FastMCP 4.x (stdio) |
Database | SQLite via aiosqlite |
Models | Pydantic v2 |
Logging | structlog (JSON / console) |
Config | YAML + Pydantic |
Tests | pytest + pytest-asyncio |
Lint | Ruff |
Container | Docker + Compose |
CI/CD | GitHub Actions |
Roadmap
DEFAIR is built MCP-first: every phase delivers the forensic capability and its MCP exposure simultaneously.
The roadmap below also closes the coverage gap with all-in-one DFIR toolboxes such as Hecatrace — without becoming one. Each tool they ship is integrated the DEFAIR way:
Wrapped, not exposed — a
BaseToolmanifest + normalizer, never a raw binary reachable through MCPStructured, not dumped — results land as artifacts, timeline events and findings in the case database, not as loose CSV/TXT files
Specialized images, not a monolith — heavy engines live in dedicated worker images (
worker-plaso,worker-malware,worker-memory, …) pulled from GHCRPinned, not
latest— every tool and rule set is version-pinned and hash-verified, and its version is recorded in eachToolRunVerified rule provenance — every YARA / Sigma source (SigmaHQ, YARA Forge and community sets) is pinned by tag or commit and hash-verified file by file; no rule is ever silently overwritten
Offline & isolated — workers run without network, with dropped capabilities and resource limits
✅ v0.1 — Core + Evidence Manager
Case & Evidence models
CLI:
case create,cases list,case get,evidence add/list/get/verifyMCP:
create_case,list_cases,get_case,register_evidence,list_evidence,get_evidence,verify_evidenceSQLite database with provenance
SHA-256 hashing + integrity verification on evidence
Structured logging with correlation IDs
Docker + CI/CD
✅ v0.1.5 — Host Wrapper + Container Orchestration
Docker SDK integration — orchestrate forensic containers from the host
One container per investigation, evidence mounted read-only
CLI:
container create/list/get/start/stop/exec/logs/removeMCP:
create_container,list_containers,get_container_info,start_container,stop_container,exec_in_container,container_logs,remove_containerPersistent workspaces at
~/.defair/workspaces/
✅ v0.2 — Windows foundation + MCP analysis
Dissect integration (host discovery, artifact identification)
13 EZ Tools with BaseTool wrappers + normalizers: MFTECmd, EvtxECmd, RECmd, PECmd, AmcacheParser, AppCompatCacheParser, JLECmd, LECmd, RBCmd, SBECmd, SrumECmd, WxTCmd, SQLECmd
Normalization layer (BaseNormalizer → unified artifact schema)
Evidence discovery with automatic tool recommendations
MCP tools:
discover_evidence,analyze_evtx,analyze_mft,analyze_registry,analyze_prefetch,analyze_amcache,analyze_shimcache,analyze_jumplist,analyze_lnk,analyze_recyclebin,analyze_shellbags,analyze_srum,analyze_wintimeline,analyze_sqliteTested on HackTheBox DFIR challenges (Jingle Bell, Recollection)
✅ v0.3 — Detection + Timeline + MCP hunting
Hayabusa v4.1 integration (4000+ Sigma rules, MITRE ATT&CK mapping)
Timeline Engine — unified timeline over all artifacts (summary, search, export CSV/JSONL)
Findings Engine — auto-created from Hayabusa high/critical detections (
FND-NNN)IOC search across all artifacts (description, data, hostname, username)
Hunting orchestration (
hunt_evtx→ detect → normalize → findings)CLI:
defair hunt,defair timeline,defair findings,defair searchMCP:
hunt_evtx,build_timeline,search_timeline,list_findings,search_ioc
✅ v0.3.1 — Prefetch analysis fix
Replaced PECmd (Windows-only) with a cross-platform Prefetch parser based on libscca
analyze_prefetchMCP tool works end-to-end in Linux containers
✅ v0.3.5 — Mass YARA + Sigma scanning
YARA mass scanner on mounted evidence (files, memory dumps, disk images)
Sigma mass scanner via Hayabusa on all EVTX sources
Default rule sets embedded in the container (YARA community rules + Hayabusa Sigma rules)
Custom rules mounting: bind-mount
/rules/yara/and/rules/sigma/for custom rulesScan results normalized as Findings with severity, confidence, and MITRE mapping
MCP:
scan_yara,scan_sigma— CLI:defair scan yara,defair scan sigma212 tests, 16 tool wrappers
✅ v0.3.6 — Hardening + reproducible images
Prerequisite: before adding more engines, make the platform match its own principles.
No shell via MCP, for real:
exec_in_containerremoved from MCP (or gated behind an explicitmcp.allow_exec: falseconfig flag, off by default) — kept in the CLI for humansrun_tool(MCP + CLI) — run a registered tool with validated arguments, recorded as aToolRun(replaces arbitrary exec for agents)Evidence allowlist —
create_containeronly mounts paths under configured evidence roots; image restricted toghcr.io/joblinours/defair*Container hardening —
cap_drop: ALL,no-new-privileges,network_mode: none, CPU / memory / PIDs limits, read-only rootfs + tmpfsDockerfile: multi-stage build (downloads in a
builderstage, no compilers in the runtime image)Pinned supply chain: EZ Tools and Hayabusa pinned by version + SHA-256; tool versions stored in every
ToolRun(rule sets: see v0.3.7)Analyst shell (CLI only):
defair container shell <case>— interactive session in the case container, evidence:ro, never exposed via MCP (≈hecatrace shell)
✅ v0.3.7 — Raijin scan engine (vendored) + verified rule sets
Raijin (Rust, YARA-X + sigma-rust) is integrated in-tree (engines/raijin/) and is the single engine behind scan_yara, scan_sigma and scan_evidence.
Engine
Raijin source vendored in
engines/raijin/as a DEFAIR-maintained fork (changes listed inengines/raijin/LICENSING.md), built from source with a pinned Rust toolchain in the image'sraijin-buildstageCold scan of the target folder only:
--lab --no-procs --scan-all-files— never live processes, never other host drivesReplaces yara-python + the unpinned Yara-Rules/rules bundle for YARA, and Hayabusa for mass Sigma scanning
Native layout detection (KAPE, Velociraptor, plain mount), original Windows path and event time kept, one detection per matched event
Raijin JSONL now carries a structured rule reference (
engine,name,id,namespace,file,tags,level) → normalizer → artifacts + findings (Sigma level / YARA score → severity, MITRE techniques from Sigma tags)
Rule sources — 13 pinned sources, ~35,000 loadable rules
Engine | Sources | Profile |
YARA | YARA Forge core |
|
YARA | YARA Forge full, Elastic protections-artifacts, ESET malware-ioc, ReversingLabs, Malpedia signator-rules, Neo23x0 signature-base, Trellix ATR |
|
Sigma | SigmaHQ core + emerging-threats add-on |
|
Sigma | SigmaHQ all rules, mdecrevoisier SIGMA-detection-rules, LOLRMM |
|
Integrity — every source pinned and verified
Sources declared in
src/defair/rules/sources.yaml, pinned insrc/defair/rules/lock/rules.lock: release tag or commit SHA (never a branch), archive SHA-256, license, and a per-file SHA-256 manifest (lock/manifests/<source>.json) — shipped inside the Python packagedefair rules lock --refreshre-pins (release archives checked against GitHub's published digest); bumping the lock is a reviewed commitdefair rules sync(image build) installs each source into/opt/defair/rules/<engine>/<source>/and verifies every file — any missing, extra or modified file aborts the buildRe-verified before every scan: a store that differs from the lock → scan refused
raijin-util validateruns per source at build (VALIDATION.txt); the build fails if a source has no loadable ruleraijin-util update/upgradedisabled — no unpinnedlatestdownloads
Collision handling — no rule silently overwritten
No basename flattening: each source keeps its upstream tree, so same-named files in different folders or sources all survive
Per-run signature tree assembled as
NN_<source>symlinks in lock order: first source wins on a duplicate Sigmaid, each YARA source loads in its own namespaceCONFLICTS.jsonper profile: identical Sigma duplicates, conflicting Sigma rules (sameid, different content — first source wins) and YARA rule names shipped by several sources (defair rules conflicts, MCPlist_rule_conflicts)YARA hits of the same rule from several sources are merged into one artifact listing every source
Provenance in results
Every finding's
detection_refs: engine, source, repo, pinned ref, rule id / name, rule file path + SHA-256, licenseCustom rules (
/rules/yara/,/rules/sigma/or--yara-rules-dir/--sigma-rules-dir) linked as99_custom, hashed into the ToolRun, taggedprovenance: custom/opt/defair/rules/NOTICElists every embedded rule set with its license (DRL 1.1, Elastic License 2.0, CC BY-SA 4.0, MIT, BSD…)CLI:
defair scan yara|sigma|evidence --profile precise|broad,defair rules status|verify|conflicts|lock|syncMCP:
scan_yara,scan_sigma,scan_evidence,get_ruleset_info,list_rule_conflictsFixed: scan findings were never created (the service read
artifact_countinstead ofartifacts_produced)Hayabusa kept for
hunt_evtxtimeline enrichment, pinned by version + SHA-256Deferred: offline rule updates from a verified bundle (
defair rules update --bundle) — rules are updated by re-pinning the lock and rebuilding the image
✅ v0.3.8 — Preprocessing & normalization pipeline
Inspired by ArtefactProcessor / PyTriage, adapted to DEFAIR's provenance model.
Two-stage normalization: tool output → normalized JSONL (
/workspace/normalized/<artifact_type>/<RUN>.jsonl, SHA-256 recorded innormalized_files) → batched bulk insert (2,000 rows per batch, numbering computed once per run instead of aCOUNT+ commit per row)Deterministic artifact ids (
uuid5(run_id, record_key)) — re-normalizing the same output yields the same ids andART-NNNnumbers, so findings stay linkeddefair normalize replay --caserebuilds a case from its JSONL (files whose hash changed are refused);defair normalize rerun RUN-xxxre-normalizes from the raw tool output after a normalizer fix;defair normalize stats RUN-xxxCommon envelope on every artifact (
provenance): tool + pinned version, run, evidence, source file, host path, channel, record id, raw value of any unparseable timestampTimestamps: ISO 8601 UTC with full source precision (EZ Tools' 7 fractional digits kept); an unparseable time stays
nullwith a reason — never replaced by "now"; ambiguousdd/mmvsmm/ddformats are refusedTimeline fields on every event:
timestamp_desc(Created ($SI), Last Executed, Event Logged…) andmessage;defair timeline export --format timesketch(MCPexport_timeline)Generic EVTX flattening:
System+EventData/UserData→ flat keys, for EvtxECmd'sPayloadand native recordsEventID knowledge base (
src/defair/data/evtx_catalog.yaml, 89 events: Security, System, Sysmon, PowerShell, RDP, Task Scheduler, Defender, WMI, BITS, USB): channel-aware type, category, description and MITRE techniques — add an event without codePure-Python fallbacks:
evtx_native(pyevtx-rs) for EvtxECmd,lnk_native(LnkParse3) for LECmd — run automatically when the tool fails, as their own ToolRun linked byfallback_ofPer-run counters in
tool_runs.normalization_stats: rows read, normalized, skipped, errors, unparseable timestamps, reasonsSchema migrations (
PRAGMA user_version): databases from earlier versions upgrade in placeMCP:
normalize_replay,normalize_rerun,get_normalization_stats,export_timelineNot copied from ArtefactProcessor: lossy
dd/mm/YYYY HH:MM:SStimestamps,datetime.now()substituted for missing times, silently swallowed exceptionsDeferred: JumpList pure-Python fallback
✅ v0.4 — Evidence sources + Orchestration + MCP profiles
🎯 Milestone: MVP MCP — one call runs a full Windows investigation, whatever the evidence format.
Evidence sources (src/defair/sources/)
Detection: KAPE (folder / VHDX), Velociraptor, FastIR, UAC, mounted filesystems, log folders, disk images (E01 / Ex01 / VMDK / VHD / VHDX / QCOW2 / raw), ZIP (plain, ZipCrypto, AES), Generaptor, DFIR-ORC — plus artifacts identified by content (magic bytes) when a collector renamed them
Collection folders registered as evidence with a tree hash (
verifydetects any added / removed / modified file)Preparation (
defair evidence prepare, MCPprepare_evidence): collections used in place, read-only; ZIP extracted (zip-slip and archive-bomb guards, ZIP-in-ZIP); Generaptor decrypted (RSA-OAEP + AES); DFIR-ORC decrypted with ANSSI's orc-decrypt (vendored inengines/orc-decrypt/, LGPL-2.1) + nested 7z; disk images carved with Dissect — no mount, no privilege, no partition offset to compute — plus a host profile artifactmanifest.jsonrecords every derived file (origin, size, SHA-256); secrets (archive password, key passphrase) travel through the exec environment, never on a command line, never stored; private keys mounted read-only at/keys(container create --keys, validated againstcontainer.key_roots)
Orchestration (src/defair/orchestrator/, src/defair/profiles/)
Declarative profiles:
windows-triage,windows-full,ransomware,persistence,registry-only,scan-onlyDAG executor: dependencies, bounded parallelism (
orchestrator.max_parallel), per-step timeout, retry with backoff, optional steps, cancellation (running tools are killed), resume (completed steps kept)Engines:
auto= EZ Tools → pure-Python fallbacks → Dissect plugins on the original image / collection;ez= no Dissect;dissect= Dissect plugins only (dissect_plugintool + generic normalizer, same artifact types as EZ Tools)Background profile runs
PRUN-NNN: detached worker, status / list / cancel / resume, dead-worker detectionrun.jsonwritten after every step and always at the end — even on failure: DEFAIR version, profile, engine, evidence preparation, each step's tools, pinned versions, fallbacks used, durations, errors, rule lockConcurrency fixes: RUN / ART / FND numbers reserved under a lock (parallel steps collided on
RUN-NNN)CLI:
defair profile list|show,defair run start --case CASE-xxx --evidence EVD-001 --profile windows-triage(≈hecatrace run -v -e),defair run status|list|cancel|resumeMCP:
prepare_evidence,list_profiles,get_profile,run_profile,analyze_evidence,get_run_status,list_runs,cancel_run,resume_runNot yet: Linux disk images (UAC collections are scanned with Raijin only), BitLocker-encrypted volumes
✅ v0.4.1 — Investigation ergonomics (from the first real E01 run)
--caseaccepts the case name everywhere (case-insensitive; ambiguous names refused), as well asCASE-YYYY-NNNor the idfindings getexplains every match: evidence file (real path in the container), event (EventID / RecordID / Computer) or YARA offset, and the exact field values / pattern that hit the rule; rule id, rule file anddefair rules show <source> <path>;--all,--json(with the raw EVTX event / hex context)artifacts get ART-NNN: every field and the full data, provenance (tool, pinned version, RUN-NNN, input), linked findings, normalized JSONL, the complete EVTX event read back from the log by record id, a hex dump around each YARA match,--context MINfor the surrounding timelineartifacts list:--contains(any field),--tool,--host,--user,--severity,--since/--until,--run,--asc, paging (--offset),--jsonContainer logs: every command (arguments with secrets redacted, output, exit code, duration), tool run (command, exit code, stderr), profile step and finding is logged as JSON to
/workspace/logs/defair.log; the container's PID 1 follows it, sodocker logs <container>shows everything. Console logs moved to stderr (stdout carries results only)container createrefuses a workspace the host user cannot write (created by DEFAIR < 0.3.6 running as root) with thechownfixHayabusa 4.x:
dfir-timelinesyntax, run from/opt/hayabusa, abbreviated levels (crit,med) mapped — critical detections now become findingsMCP:
get_artifact,get_finding,show_rule;list_artifactsgains every filter + paging
✅ v0.4.5 — Windows coverage completion
Fixes first
MFTECmd
$Joutput is normalized aswindows.usn.journal_entry(it was a file entry without timestamp);-m $MFTresolves parent paths; Dissectusnjrnlrecords use the same typeSrumECmd: one artifact type per SRUM table (app resource usage, network usage / connectivity, energy, push notifications) instead of
network_usagefor everythingMCP
analyze_srumforwardsregistry_hive; MFTECmd--bdltakes the drive letter only
Remaining EZ Tools (pinned by SHA-256 like the others)
RecentFileCacheParser (Windows 7 program execution), SumECmd (User Access Logging — who reached which Windows Server role, from where), bstrings (pattern search in any binary; runs with a pseudo-terminal on stdin), rla (dirty hive replay)
Profile steps can read another step's output:
input: step:hives_replay(+input_fallback) — rla replays the.LOG1/.LOG2of dirty hives into the workspace, RECmd parses the clean copies
NTFS depth — native parsers on dissect.ntfs, no mount, E01 / VMDK / VHDX / raw
$I30INDX slack (indx_native, INDXRipper approach): every directory of every NTFS volume;$FILE_NAMEremnants of deleted / renamed files with their four$FNtimes; entries still live (allocation or$INDEX_ROOT) are dropped$LogFile(logfile_native): file names linked / unlinked, FILE records created / freed,$FILE_NAMEcreated / removed; every other operation is counted as not decoded in the run statistics, never guessedUSN journal in
windows-triage
Extra Windows artefacts (pure-Python parsers, Dissect plugins as fallbacks)
Defender MPLog (detections, processes Defender measured, SDN file hashes, exclusions), PowerShell
ConsoleHost_history, Scheduled Task XML (actions, triggers, principal —defusedxml), WebCache (IE / legacy Edge history incl. Explorerfile://accesses, downloads, cookies), RDP bitmap cache (tiles rebuilt as PNG + a collage), IIS W3C logsLocal times without a zone (task registration, some MPLog lines) are kept as text, never assumed UTC
Typed EVTX views — src/defair/data/evtx_views.yaml, on top of the EventID catalog
logons,process_creation,services,scheduled_tasks,rdp,powershell,account_changes,log_cleared,defender,network_shares,kerberos_ntlm— filtered, column-projected views over the EVTX artifacts of any parser (EvtxECmd, native, Dissect); schema v4 indexes the EventIDCLI
defair evtx views/defair evtx view logons --case X [--since/--until/--host/--user/--event-id]
Chainsaw — second Sigma engine, pinned (2.16.5, SHA-256 = GitHub digest), on the verified DEFAIR rule store (never its own bundle): rule file, SHA-256, source, pinned ref resolved by Sigma id; findings like Raijin's — defair hunt --engine chainsaw --rule-profile precise|broad
Host profile — hostname, domain, OS build, architecture, timezone, install date, users, IPs, installed applications (Dissect, also on collections), registry time zone / network profiles / USB devices / services, computer names in the logs — every fact with its source, disagreements listed as conflicts, stored as one windows.system.host_profile artifact per evidence (≈ Hecatrace systeminfo.txt) — defair host profile, profile action host_profile
Strings + watchlists
Strings extraction (
strings_native): ASCII + UTF-16LE ofpagefile.sys/swapfile.sys, streamed through Dissect (never copied), written as a TSV index;hiberfil.sysreported as skipped (compressed — v0.8), unallocated space in v0.5Watchlists: built-in (
offensive_tools,lolbins,rmm,exfiltration) + per case (/workspace/watchlists/*.yaml), literal or regex terms; one batch search over the normalized artifacts (hits →ART-NNN), the strings index and the raw collection files (ASCII + UTF-16), ripgrep-backed (pinned) with a pure-Python fallback; JSON report + optional findings —defair watchlist list|show|searchMCP:
analyze_usn,analyze_ual,hunt_chainsaw,list_evtx_views,get_evtx_view,get_host_profile,list_watchlists,search_watchlistNot yet: hiberfil decompression (v0.8), JumpList pure-Python fallback
✅ v0.5 — Supertimeline
Worker images — heavy engines out of the main image
workers:indefair.yaml(image pinned by version tag —latestis refused — plus memory / CPU / PIDs / timeout)The case container has no Docker socket: the host starts a short-lived job container next to it (
services/worker_service.py) with the case container's evidence (read-only, same paths) and workspace mounts — never/keys— and the same hardening, network alwaysnoneThe job is
/workspace/jobs/<RUN>/job.json, written by the case container: argv lists only, never a shell;python -m defair.workers.entryruns it, writesstatus.jsonafter every step, logs to the case log (docker logsshows the job); SIGTERM kills the running tool and recordscancelledCI builds a matrix of images:
ghcr.io/joblinours/defairandghcr.io/joblinours/defair-worker-plaso, each checked after publication (python -m defair.workers.check)
worker-plaso image (docker/worker-plaso.Dockerfile)
Plaso 20260720 and all 73 Python dependencies installed with
pip --require-hashesfromdocker/worker-plaso.requirements.txtThe Sleuth Kit 4.15.0 + libewf-legacy 20140816 compiled from tarballs pinned in
docker/worker-plaso.checksums.sha256(E01 / Ex01 read directly); no compiler in the runtime image
Jobs — defair supertimeline start --case X --evidence EVD-001 --mode plaso|bodyfile|both|unallocated|all
plaso:log2timelineinto/workspace/plaso/<EVD>.plaso, thenpsort -o json_line. The storage is reused (only psort runs again) when it was built from the same evidence hash with the same parsers and time zone and still matches its recorded SHA-256; built differently → refused unless--overwrite; a half-built storage is deletedbodyfile(disk images):mmls, thenfls -r -mper partition — the filesystem type is tried explicitly (-f ntfs,fat,ext…, themmlsdescription first) instead of TSK's autodetection; the bodyfile is parsed directly (one event per distinct MACB time, likemactime, epoch = UTC)unallocated:blklsper partition streamed into the strings extractor →/workspace/strings/<EVD>/unallocated-<n>.tsv, searchable withsearch_watchlist
Timeline Engine
Schema v5:
timeline_eventstable (no ART-NNN), composite(case_id, timestamp)index, FTS5 index on message + file nameStreamed import (bounded memory — 554,138 events of a real Windows 10 E01 imported in 60 s with 173 MiB), normalized JSONL + SHA-256, deterministic ids (re-import replaces),
normalize replayrestores events too,normalize rerun RUNre-imports a supertimeline runTimestamps at source precision (Plaso
date_time: FILETIME keeps 100 ns); Plaso's "not a time" staysnullwith the raw valuetimeline summary|search|exportmerge artifacts and events in time order:--sources artifacts,plaso,tsk,--parser,--type,--offset; text search = LIKE on artifacts, FTS5 phrase on events; exports are streamed with no row limit (the former 100,000 cap is gone)run manifestper job (/workspace/jobs/<RUN>/manifest.json): spec, runner steps, worker image + digest, Plaso / TSK versions, storage reuse and hash, import countsMCP:
build_supertimeline,get_supertimeline_status(imports when the job ends),cancel_supertimeline;search_timeline/export_timelineextended to Plaso and Sleuth Kit sourcesNot a profile step: profile runs execute inside the case container, which cannot start containers — the supertimeline is started from the host (CLI / MCP)
📋 v0.6 — Reporting + REST API
Forensic reports (Markdown, HTML, JSON) — findings, host profile, timeline highlights, tool runs and versions
Case export — per-category CSV/JSONL tree for human review (Timeline Explorer, spreadsheets)
Optional export connectors: Timesketch (timeline) and OpenSearch (bulk JSONL) — the case database stays the source of truth
REST API (FastAPI + OpenAPI)
MCP:
generate_report,export_case
📋 v0.7 — Malware & document triage
Dedicated
worker-malwareimage (offline, no network)File extraction from evidence into the workspace with hash + source provenance (
extract_file)capa (capabilities), FLOSS (obfuscated strings), pefile (PE metadata), ssdeep / TLSH (fuzzy hashing)
Documents: oletools, oledump, pdfid, pdf-parser, ExifTool metadata
ClamAV with a pinned, offline signature database
YARA triage of extracted files through Raijin (same pinned rule sets as v0.3.7)
Results normalized as artifacts + findings (MITRE ATT&CK from capa)
MCP:
extract_file,triage_file,scan_clamav
📋 v0.8 — Memory forensics
Volatility 3 in a dedicated
worker-memoryimage (pinned symbol tables for offline use)Profiles:
memory-triage(pslist/pstree, cmdline, netscan, malfind, svcscan, dlllist)Memory artifacts correlated with disk artifacts in the timeline and findings
YARA + strings on memory dumps
MCP:
analyze_memory,run_profile memory-triage
📋 v0.9 — Carving + deep disk
Dedicated
worker-carvingimage: bulk_extractor (emails, URLs, IPs, credit cards…), PhotoRec / foremost / scalpel, binwalkCarved items registered as derived evidence (parent evidence + offset + hash)
Image verification:
ewfverify,hashdeepintegrated intoverify_evidenceMCP:
carve_evidence,run_bulk_extractor🎯 Milestone: coverage parity with all-in-one DFIR toolboxes — with structured results, provenance and MCP access
📋 v0.10+ — Beyond parity
Email: PST/OST/MBOX parsing, attachment extraction, OCR on attachments
Cloud & AD sources (ArtefactProcessor plugins): DFIR-O365RC, Google Workspace, ADTimeline, ADAudit
Zeek / TShark / Suricata — network DFIR
Linux DFIR — journald, SSH, cron, systemd, Docker artifacts; configurable log globs (auth, audit, nginx, apache…); Sigma on Linux logs via Raijin
Web UI
RBAC, audit trail, SBOM + image signing
📋 v1.0 — DEFAIR
Complete forensic workflow
Windows + Linux + Memory + Network
Stable MCP + API
Reproducible reports
Offline installation
Production-ready security
See defair_Roadmap.md for the full detailed roadmap.
Forensic engines
Engine | Purpose | Worker image | Phase | Status |
Dissect | Image carving without mount (v0.4), plugins as parsing fallback / |
| v0.2 / v0.4 | ✅ |
orc-decrypt (ANSSI) | DFIR-ORC archive decryption |
| v0.4 | ✅ |
EZ Tools (13) | Windows artifacts (MFT, EVTX, Registry, Amcache, LNK, SRUM, ...) |
| v0.2 | ✅ |
Hayabusa | EVTX hunting timeline (hayabusa-rules, distinct provenance) |
| v0.3 | ✅ |
YARA (yara-python) | File pattern matching — replaced by Raijin in v0.3.7 | — | v0.3.5 | ⛔ |
Raijin (vendored) | YARA-X + Sigma cold scanner, 13 pinned rule sources |
| v0.3.7 | ✅ |
EZ Tools (remaining) | RecentFileCacheParser, SumECmd, bstrings, rla |
| v0.4.5 | ✅ |
NTFS native (dissect.ntfs) |
|
| v0.4.5 | ✅ |
Chainsaw | Second Sigma engine on the pinned rule store |
| v0.4.5 | ✅ |
ripgrep | Watchlist / IOC batch search |
| v0.4.5 | ✅ |
Plaso | Supertimeline, multi-source timestamp normalization |
| v0.5 | ✅ |
Sleuth Kit | Bodyfile filesystem timeline, unallocated space ( |
| v0.5 | ✅ |
capa / FLOSS / pefile | Malware capabilities, obfuscated strings, PE metadata |
| v0.7 | 📋 |
oletools / Didier Stevens suite / ExifTool | Office, PDF and metadata triage |
| v0.7 | 📋 |
ClamAV | AV scanning (pinned offline signatures) |
| v0.7 | 📋 |
Volatility 3 | Memory forensics (processes, network, DLLs, persistence) |
| v0.8 | 📋 |
bulk_extractor / PhotoRec / foremost / scalpel | Feature extraction and file carving |
| v0.9 | 📋 |
Zeek / TShark / Suricata | Network traffic analysis, IDS |
| v0.10+ | 📋 |
Out of scope: acquisition tools (dc3dd, imaging) — DEFAIR analyses evidence, it does not acquire it.
Development
# Run tests
pytest tests/ -v
# Lint
ruff check src/ tests/
# Run CLI
defair --help
# Run MCP server
defair-mcpAdding a new MCP tool
Add the service function in
src/defair/services/Add the CLI command in
src/defair/cli/Add the MCP tool in
src/defair/mcp_server/server.pywith@mcp.tool()Add tests (unit + CLI + MCP integration)
Both CLI and MCP must call the same service function
License
MIT
DEFAIR: a reproducible DFIR platform capable of orchestrating forensic engines, unifying their results, and exposing the investigation to a human or an agent via CLI, API, and MCP.
Available Tools
7 toolscreate_caseCreate CaseC
Create a new forensic investigation case.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name of the case (e.g. "Incident host-01 ransomware"). | |
| description | No | Optional longer description of the case. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it says nothing beyond the act of creation: no persistence guarantees, no required auth/permission level, no indication of whether the case is immediately visible to get_case/list_cases, and no side effects. For a mutation tool this is a meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the verb-resource pair front-loaded and zero filler. It is efficient, though its brevity is partly under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the schema fully covers the two simple parameters. However, as an unannotated mutation tool it omits permissions, side effects, and any post-creation workflow hints, leaving the definition minimally but not fully adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents both 'name' and 'description' including an example and a default. The description adds no parameter meaning beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a ... case') and the domain ('forensic investigation'), which is enough to distinguish it from the read-oriented siblings (list_cases, get_case, verify_evidence). It does not explicitly name or contrast the nearest alternative, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to create a case versus reusing an existing one, no prerequisites (e.g., authentication or permissions), and no mention of sibling tools such as register_evidence that would follow. The agent must infer context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_caseGet CaseB
Get details of a specific forensic case.
| Name | Required | Description | Default |
|---|---|---|---|
| case_id | Yes | The case number (e.g. "CASE-2026-001") or internal UUID. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Get' weakly implies a read, but the description never states that this is a non-mutating lookup, what permissions are required, or whether case data is sensitive/auditable. With zero annotation coverage this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no waste, front-loaded with the verb and resource. Appropriate size for a one-parameter lookup tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be described, and the single parameter is fully documented in the schema. But with no annotations and no usage guidance, an agent still lacks context on safety profile and when to prefer this over list_cases/get_evidence. Minimum viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the case_id description already provides format examples (CASE-2026-001, internal UUID). The description adds nothing about the parameter, so the baseline of 3 applies – the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('details of a specific forensic case'), clear enough to distinguish from siblings like list_cases (plural, listing) and create_case (mutation). However it doesn't explicitly contrast with get_evidence or explain what 'details' includes, leaving some ambiguity about sibling selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, no mention of alternatives. The singular 'a specific forensic case' weakly implies you need a known case_id, but nothing tells the agent when to reach for this versus list_cases or get_evidence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evidenceGet EvidenceC
Get details of a specific evidence item.
| Name | Required | Description | Default |
|---|---|---|---|
| evidence_id | Yes | Evidence number (e.g. "EVD-001") or internal UUID. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden but only implies a read via 'Get'. It says nothing about whether the evidence must exist, error behavior for unknown IDs, any access restrictions, or what happens with the two accepted ID formats. Minimal disclosure for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words and the verb/resource established immediately. It is efficient, though almost too terse to carry any extra value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and there is only one well-documented parameter. The definition is adequate for a simple lookup but omits error/edge behavior and sibling differentiation that would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single parameter is documented with concrete examples ("EVD-001" or internal UUID). The description adds no additional meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('a specific evidence item'), which is clear and unambiguous. However, it does nothing to distinguish itself from siblings like get_case or list_evidence/verify_evidence, leaving the agent to infer from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus list_evidence, verify_evidence, or register_evidence. The description offers no context, prerequisites, or exclusions, so an agent must guess that this is the single-item lookup counterpart to list_evidence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_casesList CasesB
List all forensic cases.
Returns a list of cases with their case number, name, status, and creation date. Cases are ordered by creation date (newest first).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does add real behavioral context: the ordering (newest first) and the fields returned. However, it is silent on whether the full result set is returned unbounded, whether pagination or limits exist, and whether any permissions are required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the purpose. The second sentence largely restates the output schema's fields, which is mild redundancy but not enough to hurt readability or comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained in prose, and the empty input schema means there are no parameter gaps to cover. The only missing piece is whether results are bounded or paginated, which is a minor gap for a zero-parameter list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate beyond confirming the operation takes no input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List all forensic cases') and the scope ('all') implicitly distinguishes it from get_case, which retrieves a single case. It is unambiguous but never explicitly contrasts itself with the sibling read tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus get_case, and no mention of whether the result set is filtered, paginated, or capped. The agent must infer that 'all' means no filtering and no pagination control exists given the empty parameter schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_evidenceList EvidenceB
List registered evidence items.
| Name | Required | Description | Default |
|---|---|---|---|
| case_id | No | Optional — filter by case number (e.g. "CASE-2026-001") or UUID. If omitted, returns all evidence across all cases. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It doesn't state whether this is a read-only operation, whether it paginates, what the return volume might be, or any auth requirements. For a list tool with no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is concise and front-loaded, but it's so terse that it sacrifices useful context. It earns points for zero waste, but doesn't provide enough substance to be considered well-structured for a list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained. The single optional parameter is fully documented in the schema. However, with no annotations and no behavioral context, the description is barely adequate for an agent to understand when and how to use this tool versus siblings like get_evidence or list_cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — the case_id parameter is fully documented in the schema including its optional nature, accepted formats (case number or UUID), and default behavior. The description adds no parameter information beyond what's already in the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource: listing registered evidence items. This distinguishes it from get_evidence (singular retrieval) and register_evidence (creation). However, it doesn't explicitly contrast with list_cases, the closest sibling, leaving some sibling differentiation implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (you call it to see a list of evidence), but provides no explicit when-to-use guidance, no exclusions, and doesn't name alternatives like get_evidence for single-item retrieval. The optional case_id parameter's behavior is documented in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_evidenceRegister EvidenceA
Register a new piece of evidence in a forensic case.
Computes SHA-256 hash of the file and stores metadata. The original file is never modified or copied.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Absolute path to the evidence file. | |
| case_id | Yes | Case number (e.g. "CASE-2026-001") or UUID. | |
| evidence_type | No | Type of evidence — one of: disk_image, memory_dump, logs, triage_archive, pcap, other. | other |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and discloses important side effects: it computes a SHA-256 hash, stores metadata, and explicitly states the original file is never modified or copied. It still omits permissions, duplicate-registration behavior, and error conditions, but the core mutation and file-safety behavior are clearly conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with the purpose front-loaded, followed by concise behavioral details. Every sentence adds useful information without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, complete parameter descriptions, and an existing output schema, the description covers the essential action and key side effects. It is nearly complete, though it could mention prerequisites such as whether the case must already exist or what happens on duplicate registration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters, including case_id format, path, and evidence_type enum values. The description adds no parameter-level meaning beyond what the schema provides, which matches the baseline of 3 for fully covered schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Register') and resource ('a new piece of evidence in a forensic case'), making the action immediately identifiable. It distinguishes itself from sibling tools like get_evidence, list_evidence, and verify_evidence by focusing on registration of new evidence rather than retrieval or verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use guidance, prerequisites, or alternative tools. It implies registration is for adding new evidence, but does not say when to prefer it over verify_evidence or how to handle existing evidence, leaving usage conditions entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_evidenceVerify EvidenceA
Verify evidence integrity by re-computing its SHA-256 hash.
Compares the current file hash against the hash stored at registration. This is a critical forensic operation to detect evidence tampering.
| Name | Required | Description | Default |
|---|---|---|---|
| evidence_id | Yes | Evidence number (e.g. "EVD-001") or internal UUID. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the core behavior (recompute hash, compare to the value stored at registration) and implies a non-destructive read, but never states permissions, whether the result is persisted or logged, or what happens on a mismatch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the mechanism front-loaded and no filler. The phrase "critical forensic operation" is mildly editorial but harmless, so structure is strong though not perfect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return shape need not be explained, and the one parameter is fully covered. The description is sufficient to call the tool, though it leaves the mismatch behavior and read/write semantics implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single evidence_id parameter is already fully documented in the schema (evidence number or UUID). The prose adds no parameter-level detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource (verify evidence) and even names the mechanism (re-computing its SHA-256 hash) plus the goal (detect tampering). This clearly distinguishes it from the read-only siblings get_evidence and list_evidence, which retrieve rather than validate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage ("critical forensic operation to detect evidence tampering"), which tells an agent the context in which it matters. However, it names no alternatives or preconditions, e.g. whether get_evidence must be called first or when verification is unnecessary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
create_case - First observed
get_case - First observed
get_evidence - First observed
list_cases - First observed
list_evidence - First observed
register_evidence - First observed
verify_evidence
TDQS
Scored across 7 tools
Each tool targets a distinct resource (case vs. evidence) and action (list, create, get, register, verify). The purposes are clear and non-overlapping, with no ambiguity in selection.
All tool names follow a consistent verb_noun pattern: list_cases, create_case, get_case, register_evidence, list_evidence, get_evidence, verify_evidence. This is predictable and easy to parse.
Seven tools are well-scoped for a forensic case management server, covering essential operations without redundancy. Each tool earns its place in the set.
The surface covers core case and evidence lifecycle operations, including creation, retrieval, listing, and integrity verification. However, it lacks update and delete operations for cases and evidence, which are common CRUD needs that could be required for full lifecycle management.
Maintenance
Related MCP Connectors
Inspection, steganography and forensics API for files and images. Bitcoin pay-per-use.
Authenticated public evidence search, verification, research jobs, exports, and webhooks.
Tamper-evident proof creation and verification for AI agents via MCP, A2A, and REST.
Zero-install remote MCP server for proof-of-existence file attestation.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceEnables EHR Copilot operations such as order cloning, queue building, and execution trace analysis over stdio.-
- FlicenseAqualityCmaintenanceEnables local forensic analysis of files by orchestrating system binaries (file, exiftool, strings, Volatility) via safe subprocess execution.3-
- FlicenseAqualityCmaintenanceMCP server for read-only forensic analysis of evidence files using local utilities (file, ExifTool, strings, Volatility).3-
- AlicenseNot gradedqualityAmaintenanceEnables agents and applications to record, query, verify, and receive hash-chained event provenance over stdio, supporting self-metering and independent audit of autonomous system activity.41 PyPI6MIT