headcleaner
headcleaner
Walk a folder of mixed documents, get back clean Markdown you can search, cite, and trust.
headcleaner is a Python command-line tool that reads a folder of documents — Word, Excel, PowerPoint, PDF, HTML, plain text, email — and turns each one into a Markdown file with a citation back to its original source. The output is portable, searchable, and ready for indexing, archiving, or handing to an AI coding assistant. Headcleaner never silently rewrites your files, never claims the output has been human-reviewed when it has not, and never talks to a network service unless you explicitly tell it to.
Why headcleaner exists
Most teams sit on folders full of .docx, .xlsx, .pdf, .html, and .eml files that are useful but hard to search, hard to cite, and hard to feed to modern AI tooling. Headcleaner is the missing layer between those folders and the systems that want to consume them: a local-first converter that produces a durable, citation-aware Markdown output you can rebuild, audit, and trust.
The core promise is small enough to fit on a sticky note: read once, write Markdown with a source citation, and never lie about whether a human has reviewed it.
What headcleaner does
Headcleaner converts a folder of mixed documents into a clean output folder you control. The same input folder produces the same output bytes every time, and every output file carries the SHA-256 hash of the source it came from.
The conversion pipeline reads source files, normalizes them through format-specific adapters, and writes either plain Markdown, an OKF v0.2 knowledge bundle, or both side by side. On top of that, headcleaner can build a local SQLite search index, generate a knowledge graph of how your documents relate to each other, and surface cited chunks to compatible AI assistants through the Model Context Protocol.
What headcleaner does not do
This list is short on purpose. Every item is a design choice, not an oversight.
It does not silently rewrite your source files. Source files are read-only inputs to headcleaner; output goes to a directory you name.
It does not claim that auto-converted output has been human-reviewed. Every emitted file starts in an explicit "human has not read this" state; changing that requires an explicit review action.
It does not install tools you did not ask for. Optional converters like OfficeCLI, LibreOffice, and Tesseract are checked for at runtime; headcleaner tells you when one is missing instead of trying to install it.
It does not talk to a network service by default. Embedding providers, remote vector databases, and MCP integration all require an explicit configuration step.
It does not rewrite your git history, publish packages, or push to a remote. Version-control operations require explicit invocation.
Three-step quick start
This is the smallest path from "I just installed headcleaner" to "I see something useful."
Step 1 — Convert a folder
Pick any folder that contains documents headcleaner can read. For a first run, a folder with one PDF, one Word file, and one HTML file is ideal.
uv run --no-sync --python 3.13 headcleaner convert ./my-folder --output ./my-folder.cleanThe convert command reads ./my-folder, normalizes every supported document it finds, and writes the results to ./my-folder.clean. By default you get both plain Markdown and an OKF v0.2 bundle side by side.
Step 2 — Look at what was produced
Open ./my-folder.clean in your file browser. You will see three things: a manifest.json that summarizes the run, an _md/ folder containing one Markdown file per source, and an okf/ folder containing the OKF v0.2 bundle with an index.md and one concept file per source.
./my-folder.clean/
├── manifest.json # run summary: what was processed, how, with what status
├── REPORT.md # human-readable run report
├── _md/ # plain Markdown, one file per source
│ ├── notes.docx.md
│ ├── report.pdf.md
│ └── page.html.md
└── okf/ # OKF v0.2 bundle, one concept per source
├── index.md # auto-generated directory index
├── notes.docx.md
├── report.pdf.md
└── page.html.mdEach generated file starts with a YAML block that names the source it came from, the SHA-256 hash of that source, the date the source was generated, and the trust state. That block is how headcleaner keeps its promise that you can always answer "where did this text come from."
Step 3 — Open the report and the manifest
./my-folder.clean/REPORT.md is a short Markdown file you can read in any editor. It tells you how many files were processed, which engine handled each one, and whether anything failed or was skipped. ./my-folder.clean/manifest.json is the same information in a structured form that other tools can consume.
What this means: if your input folder had twelve documents and your run produced twelve Markdown files plus an OKF bundle plus a manifest, the conversion is healthy. If the report shows files in the skipped or failed state, jump to the Troubleshooting guide — those states almost always mean an optional tool is missing, not that your project is broken.
A simple visual
The flow is small enough to draw:
Source folder on the left, the headcleaner pipeline in the middle, output folder on the right. The four purple cards underneath are the rebuildable derivatives that fall out of the pipeline. The cyan card at the bottom is the local SQLite search index, built from the cited chunks.
What to read next
Pick the path that matches what you came here to do.
I have never used headcleaner and want to install it. Start with the Installation guide, then walk through the First run guide. Both are written for someone who has never run a Python CLI tool before.
I want to understand what each command does. Go to the CLI reference, organized by what you are trying to accomplish.
I want to use headcleaner with an AI coding assistant. Read Working with AI assistants and then MCP client setup.
I want to add headcleaner to a CI pipeline. Start with CI integration and the tutorial on CI integration.
I want to extend headcleaner with a new file format or tool. Go to the Contributor onboarding and then the Tool and engine development guide.
I am developing or committing a change. Follow the development workflow and the documentation governance.
I want to understand the safety and trust model before I commit to using headcleaner. Read the Safety overview.
Documentation map by reader goal
The complete documentation is organized by reader, not by source module. Each path below is a coherent walk that answers a specific question.
If you want to… | Read |
Install headcleaner on Windows, macOS, or Linux | |
Run your first conversion and understand the output | |
Understand the terms OKF, citation, FTS5, and trust | |
Build the everyday workflow that fits how I actually work | |
Read the output files and the report | |
Know whether the output is trustworthy | |
Set up local search and graph over the output | |
Use headcleaner with a coding assistant | |
Debug a skipped check, missing tool, or wrong exit code | |
Look up a specific command, flag, or behavior | |
Understand a specific engine, install hint, or skip behavior | |
Configure headcleaner with a project settings file | |
Add a new adapter, engine, or configuration field | |
Read the architecture and the canonical data model | |
Understand the trust and safety guarantees | |
Plan, implement, audit docs, and prepare a commit |
Phase R10 agentic workspace
R10 adds Stack-confined Wiki and Memory repositories, a separate evidence-bearing
function graph, seven closed read-only agent tools, cited chat operations, a shared
dashboard/TUI snapshot, and deterministic offline agentic evaluation. The strict
commands live under headcleaner stack; see the
agentic workspace guide,
CLI reference, and
MCP migration reference.
Human review remains journal-only: generated Wiki content is proposal-only, converted
document bytes never acquire reviewed status automatically, and only canonical human
decisions in .headcleaner/review/decisions.jsonl confer review authority. Hosted
execution requires explicit policy, consent, budgets, and credential resolution through
the authorized egress boundary. Ollama, LM Studio, and generic loopback execution have
zero hosted egress. Artifacts remain below the selected Stack root; events and errors are
redacted; --dry-run and --no-store suppress durable effects; and evaluation preserves
unknown cost as unknown rather than zero.
R10 technical completion is not a public release. It does not establish clean-install, distribution, operating-system, CI, packaging, publication, or release evidence; those remain governed by the release guide.
License
Apache-2.0. See LICENSE.
Phase 4 local automation
Phase 4 adds cited context retrieval, deterministic profiles, explicit local jobs, connector preview/apply, plugin inspection, and versioned JSON events. These are local-first: convert starts no service, no command promotes human review, connector deletion is unavailable, and outbound delivery is disabled without an explicit endpoint configuration. See Automation API, Plugin contract, and Connector synchronization.
Phase R7 human review authority
Phase R7 formalizes human review as strictly journal-only (ADR 0015):
Human review decisions are appended exclusively to
<bundle>/.headcleaner/review/decisions.jsonl.Converted markdown files and frontmatters are 100% immutable and never modified by review commands.
Review status is dynamically derived via
read_review_projection(bundle_root).Reversals append a superseding record linking back to prior decision IDs.
Direct frontmatter mutation requests are refused with
REVIEW_LEGACY_MUTATION_REFUSED.