datamining-skill
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@datamining-skillProfile events.csv and extract all email addresses into emails.csv."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DataMining Skill
Mine huge files on a small machine, privately, and let an AI assistant drive it.
A local-only, streaming-first data mining toolkit. It profiles, partitions and mines CSV, JSONL and log files far larger than your RAM in constant memory, survives being killed at any instant without losing or duplicating a single record, and plugs into Claude Code, Claude Desktop, Cursor and any other Model Context Protocol client as a tool.
Private by design. Everything runs on your machine. No network code, no telemetry, no third-party packages. Your data is never sent anywhere.
Constant memory. A 100 GB file on an 8 GB laptop is fine: every buffer is bounded, so memory depends on configuration, never on file size.
Crash-safe and resumable. Kill it, lose power, run it again: the result is exactly what an uninterrupted run would have produced.
AI-ready. A built-in MCP server exposes
profile_datasetandmine_datasetto your assistant, with streamed progress and strict, workspace-confined file access.
Contents
Quick start · Use it from an AI assistant (MCP) · Command line · Python API · How it works · Guarantees · Limitations · Development · Contributing
Related MCP server: Aleph
Quick start
Requires Python 3.11 or newer. From a clone of this repository:
pip install .Inspect a file without reading it into memory, then mine it:
# What is in this file? (milliseconds, constant memory)
datamining-skill profile events.csv
# Extract every e-mail address into a result file (resumable, crash-safe)
datamining-skill mine events.csv emails.csvprofile prints a JSON description (size, encoding, format, columns or keys, estimated
record count). mine writes emails.csv and prints a JSON summary. If mine is
interrupted, run the same command again: it resumes where it stopped.
Measured on a synthetic 10 GiB CSV (Windows development machine):
Operation | Time | Memory growth |
Profile the file | ~50 ms | ~0.6 MiB |
Plan chunks for 2 GiB of free RAM (34 chunks of 307 MiB) | ~26 ms | ~200 KiB |
Use it from an AI assistant (MCP)
datamining-skill mcp runs a Model Context Protocol server
on standard input/output, so assistants can inspect and mine local files on your behalf.
Tool | What it does |
| Size, encoding, format, columns or JSON keys and an estimated record count for a CSV, TSV, JSONL or log file. Read-only. |
| Extracts data (e-mail addresses by default) into a |
You decide which directories the assistant may touch. Every file argument is resolved
(symlinks and .. included) and must lie inside a directory you list with --allow-dir;
anything else is refused. The first --allow-dir is the workspace for relative paths.
Claude Code
claude mcp add --scope project datamining-skill -- \
datamining-skill mcp --allow-dir /absolute/path/to/your/dataor add the equivalent entry to .mcp.json at the project root:
{
"mcpServers": {
"datamining-skill": {
"command": "datamining-skill",
"args": ["mcp", "--allow-dir", "/absolute/path/to/your/data"]
}
}
}Run /mcp inside Claude Code to see that the server is connected, then ask, for example:
"Profile events.csv, then extract all e-mail addresses from it into emails.csv."
Claude Desktop
Edit claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/, Windows:
%APPDATA%\Claude\) and restart the app:
{
"mcpServers": {
"datamining-skill": {
"command": "datamining-skill",
"args": ["mcp", "--allow-dir", "/absolute/path/to/your/data"]
}
}
}Cursor
Create .cursor/mcp.json in your project (or ~/.cursor/mcp.json for all projects):
{
"mcpServers": {
"datamining-skill": {
"command": "datamining-skill",
"args": ["mcp", "--allow-dir", "${workspaceFolder}"]
}
}
}Configuration notes
Use the full path to the command if it is not on the client's
PATH, for example when you installed into a virtual environment. Either point at the console script (/path/to/.venv/bin/datamining-skill, or...\.venv\Scripts\datamining-skill.exeon Windows) or at the interpreter:"command": "/path/to/.venv/bin/python", "args": ["-m", "datamining_skill", "mcp", "--allow-dir", "..."].Windows paths in JSON need doubled backslashes (
"C:\\Users\\you\\data"); forward slashes also work.You can list several
--allow-diroptions, or setDATAMINING_SKILL_ALLOWED_DIRS(paths separated by;on Windows,:elsewhere). With neither, the server uses its working directory, and refuses to if that is the filesystem root.Results, state and scratch files go to
<first allowed dir>/.scratch/and the result path you name. Add.scratch/to your.gitignore.Verify an installation without a client:
python scripts/mcp_smoke_test.pyperforms a full handshake, profile and mine against the installed command and exits non-zero on any problem.
Protocol and security
The server is written against the standard library only and speaks both generations of MCP:
the stateful initialize handshake (revisions 2024-11-05 to 2025-11-25) and the stateless
2026-07-28 revision (server/discover, per-request metadata). Messages are newline-delimited
JSON-RPC 2.0; nothing but protocol messages is ever written to stdout (logs go to stderr).
Because tool arguments come from a language model, treat them as untrusted:
Paths are confined to the allowed directories; relative paths resolve inside the workspace. Symlinks and junctions are followed before the check, UNC and device paths, alternate data streams, reserved device names and control characters are rejected outright, and nothing is URL-decoded or expanded (
%00,~and$HOMEare ordinary file-name characters).mine_datasetonly writes.csv,.jsonlor.ndjsonfiles, never overwrites an existing result unlessoverwriteis set, and refuses to write over the source.Custom regular expressions are off by default: a model-written pattern can be crafted to backtrack catastrophically and stall your machine. Start the server with
--allow-custom-patternsonly if you accept that. The built-in e-mail extractor is safe on hostile input.
See docs/privacy-and-security.md for the full threat model.
Command line
datamining-skill profile PATH [--config FILE]
datamining-skill mine SOURCE OUTPUT [--workspace DIR] [--pattern REGEX [--fields a,b]]
[--overwrite] [--csv-formula-guard]
datamining-skill mcp [--allow-dir DIR ...] [--allow-custom-patterns]All commands accept --log-level {DEBUG,INFO,WARNING,ERROR}, --no-logs and --debug.
Structured JSON logs go to stderr; results go to stdout. python -m datamining_skill.cli ... is
equivalent.
A failure prints a single error: line, never a Python traceback; add --debug to see the
traceback as well.
Exit code | Meaning |
0 | Success |
1 | Unexpected internal error (rerun with |
2 | Unsupported or unrecognised format (empty, binary, compressed, UTF-16/32 for |
3 | Unavailable source, invalid configuration, file-system error, low memory, job already running, or another library error |
4 |
|
130 | Interrupted |
141 | The reader of stdout closed the pipe (for example ` |
Mining examples:
# JSONL output instead of CSV (chosen by the extension)
datamining-skill mine events.csv emails.jsonl
# Extract your own pattern: each capture group becomes a column
datamining-skill mine app.log errors.csv --pattern "(\d{4}-\d\d-\d\d) ERROR (\w+)" --fields date,code
# Start over, discarding earlier progress and replacing the output
datamining-skill mine events.csv emails.csv --overwrite
# Opening the result in a spreadsheet? Neutralise formula injection
datamining-skill mine events.csv emails.csv --csv-formula-guardState and scratch files live in ./.scratch/ (override with --workspace). A fresh run refuses
to overwrite an existing output file that it did not create.
Python API
from datamining_skill import create_profiler, run_mining
# Profile: size, encoding, format, structure, estimated record count
profile = create_profiler().profile("events.jsonl")
print(profile.data_format, profile.records.count, profile.structure.fields)
# Mine with managed state: resumes automatically if interrupted
summary = run_mining("events.csv", "emails.csv", on_progress=lambda p: print(p))
print(summary.to_dict())Each stage is also available on its own:
from datamining_skill import StateManager, create_chunking_engine, create_orchestrator
profile = create_profiler().profile("events.csv")
# 1. Plan: memory-aware, newline-aligned byte ranges (a lazy generator; nothing is read)
for chunk in create_chunking_engine().plan("events.csv", profile):
print(chunk.chunk_id, chunk.start_byte, chunk.end_byte)
# 2. Mine with your own state database
with StateManager("data/run-001.sqlite3") as state:
orchestrator = create_orchestrator(output_path="data/results.csv", state=state)
summary = orchestrator.run("events.csv") # safe to re-run after a crashWrite your own mining logic
An ExtractionStrategy is a pure function from one line to zero or more records. It performs
no I/O; reading, writing and checkpointing are handled for you.
from collections.abc import Iterator, Sequence
from datamining_skill import StateManager, create_orchestrator
class StatusCodes:
fields = ("status",)
def extract(self, line: str) -> Iterator[Sequence[str]]:
if " 500 " in line:
yield ("server-error",)
with StateManager("data/run-002.sqlite3") as state:
create_orchestrator(
output_path="data/errors.jsonl", state=state, strategy=StatusCodes()
).run("access.log")How it works
flowchart LR
F[("Huge file")] --> P["Data Profiler<br/>size, encoding, format,<br/>structure, record estimate"]
P --> C["Chunking Engine<br/>newline-aligned byte ranges<br/>sized from free RAM"]
C --> S[("State Manager<br/>SQLite, WAL")]
S -->|"claim next PENDING chunk"| W["Miner Worker<br/>seek, read line by line,<br/>apply strategy"]
W --> T["chunk_N.tmp"]
T --> A["Result Aggregator<br/>truncate to last commit,<br/>append, fsync"]
A --> O[("results.csv / .jsonl")]
A -->|"COMPLETED + output length<br/>in one transaction"| SStage | What it guarantees |
Data Profiler | Reports size, encoding, format and structure from bounded samples; the record count is extrapolated from windows spread across the file. Memory is O(1). |
Chunking Engine | Chunk size is |
State Manager | Every chunk is PENDING, IN_PROGRESS, COMPLETED or FAILED in a SQLite database (WAL mode). Chunks left IN_PROGRESS by a crash return to PENDING automatically. |
Miner Worker | Reads exactly |
Result Aggregator | Appends a finished chunk to the output idempotently: the output's committed length is recorded in the same transaction as |
The code follows Clean Architecture; dependencies point inward only:
flowchart TB
CLI["cli.py and bootstrap.py<br/>composition root and delivery"] --> I
I["infrastructure (adapters)<br/>file streaming, SQLite state, scratch and output stores,<br/>MCP server, JSON logging"] --> A
A["application (use cases)<br/>profiler, chunking engine, worker and strategies,<br/>aggregator, orchestrator"] --> D
D["domain<br/>models, ports (Protocols), exceptions"]Design notes, diagrams and the crash-recovery analysis are in docs/architecture.md.
Guarantees
Local only. The runtime imports no networking modules and has no dependencies.
Source files are read-only. They are never modified, copied or indexed.
Footprint you can see. Besides the result file you name, only a metadata-only state database and short-lived per-chunk scratch files are written, confined to
.scratch/ordata/(never the system temporary directory), owner-only (0600/0700on POSIX, a protected ACL on Windows), removed as chunks merge.No content in logs or errors. Logs and messages carry names, sizes, ids and counts.
Data is data. Input is tokenised, never evaluated: no
eval,pickleor dynamic import on the data path; JSON, CSV and regular-expression parsing use bounded, standard-library tools.Exactly-once results. Verified by killing real processes at three different points while mining a 50 MiB file: the output matched the planted records exactly, in order, every time.
Limitations
Compressed (
.gz,.zst, ...) and binary formats (Parquet, SQLite) are rejected with a clear error; decompress or convert them first.A "record" is a physical line. Quoted CSV fields or log entries that span lines can be split across chunk boundaries.
Mining supports ASCII-compatible encodings (ASCII, UTF-8, cp1252, ISO-8859-1). UTF-16/32 can be profiled but not mined.
Mining is single-threaded. The CLI,
run_miningand the MCP tool hold a cross-process lock per job, so a second process on the same source and output is refused ("already running"). Code that drivescreate_orchestratordirectly must provide its own single-writer guarantee.If the source is truncated or replaced while it is being mined, the affected chunks fail with an explicit reason instead of silently dropping records. A file that is replaced by another of the same size cannot be detected.
A failed chunk that succeeds on a retry is appended after the others, so output order can then differ from file order (no record is lost or duplicated).
The MCP server handles one request at a time. It cannot interrupt a running tool call, but you can stop the process safely at any moment and call
mine_datasetagain to resume.Power loss can lose the most recent checkpoints (never corrupt data), so a few chunks may be mined again; this is why extraction must be idempotent, which it is by construction.
Development
git clone <your fork> && cd datamining-skill
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
python -m mypy # strict type checking
python -m pytest # full suite: crash recovery, memory bounds, path security, MCP protocolGitHub Actions runs the type check and the tests on Linux, macOS and Windows with Python 3.11 to 3.14, then builds the package and smoke-tests the installed wheel in a clean environment. Validation scripts you can run yourself:
python scripts/generate_dummy_csv.py --size-gb 2 --output data/dummy_2gb.csv # memory check
python scripts/simulate_crash_resume.py # 100 chunks, killed at 50, resumed
python scripts/simulate_mining_crash.py # 50 MiB mined, killed at chunk 3, three crash points
python scripts/mcp_smoke_test.py # MCP handshake and tools against the installed commandProject layout:
src/datamining_skill/
domain/ value objects, exceptions, ports (no dependencies)
application/ profiler, chunking engine, worker, strategies, aggregator, orchestrator
infrastructure/ streaming, encoding, SQLite state, scratch/output stores, MCP server, logging
bootstrap.py composition root (create_profiler, run_mining, create_mcp_server, ...)
cli.py command line: profile, mine, mcp
tests/ pytest suite scripts/ validation scripts
docs/ architecture and security notesContributing
Contributions are welcome under any name or pseudonym. Please read CONTRIBUTING.md and the Code of Conduct. Report security issues privately as described in SECURITY.md. Changes are recorded in CHANGELOG.md.
License
MIT. See LICENSE. Copyright (c) 2026 DataMining Skill Contributors.
This server cannot be deployed
Maintenance
Related MCP Connectors
Open, inspect, filter, edit and convert xlsx and csv files from your AI chat. Processing is local.
Open, inspect, filter, edit and convert xlsx and csv files from your AI chat. Processing is local.
Open, inspect, filter, edit and convert xlsx and csv files from your AI chat. Processing is local.
Use your Mac, Windows or Linux computer from ChatGPT, Claude or Codex: files, commands, documents.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceEnables AI assistants to intelligently search and explore local file systems using native Unix commands (ripgrep, find, ls) with token-optimized output, automatic pagination, and multi-layer security validation.164 npm2-
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to analyze documents larger than their context window by loading files into RAM and querying them via search, navigation, and Python execution tools. Supports recursive reasoning to process massive datasets in chunks using sub-agents.218MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to interact with local CSV and Parquet data files through natural language queries, facilitating tasks like summarizing datasets or retrieving specific information.4-
- FlicenseAqualityDmaintenanceEnables Claude to analyze local CSV or Parquet files, handling larger datasets without uploading full files.318-