Skip to main content
Glama
BenyD

Haypile

Haypile

Private document search, for you and your agents. One binary that watches your folders, indexes every document, and hands your agent the right passages over MCP, each with a file and page citation. Nothing ever leaves your machine.

Everyone says finding information in your files is like finding a needle in a haystack. Haypile is the haystack that finds its own needles.

hay demo: add a folder, search by meaning, verify zero outbound connections

brew install BenyD/tap/hay
hay add ~/Documents
claude mcp add --transport http haypile http://localhost:11500/mcp

That is the whole setup. Claude Code (or Cursor, or anything that speaks MCP) can now search everything you indexed: search_documents returns ranked passages with citations, and the agent answers from them instead of guessing.

It is a full standalone CLI too:

hay search "termination clause"
hay ask "what did the Meridian contract say about termination?"

Search understands meaning, not just words: "agreement cancellation" finds termination clauses. Exact identifiers still match exactly. Every result cites its source file and page.

Install

brew install BenyD/tap/hay

Or grab a binary from releases, or run the install script:

curl -fsSL haypile.sh | sh

On Windows, run this in PowerShell:

irm https://haypile.sh/install.ps1 | iex

One binary, around 56MB, with the embedding model inside. No Python, no Docker, no vector database, no model downloads, no network. Text documents index fully offline; scanned PDFs need a local vision model for OCR (hay llm setup).

Related MCP server: Markdown RAG Documentation

Why

An agent can open any file you point it at. Finding the right one is the problem: grep matches words, not meaning, and a question about termination clauses does not contain the words the contract used. Reading whole documents to find one passage spends the context window the task needed. Cloud document tools fix this by uploading everything to someone else's computer, and self-hosted RAG stacks fix it with Python environments, Docker, and a vector database to babysit.

Haypile is the missing option: one binary, point it at folders, done.

  • Agent-ready. MCP and REST on localhost:11500. search_documents gives Claude Code, Cursor, or your own scripts cited passages from your documents.

  • Hybrid search. Semantic and keyword, merged. Paraphrases match by meaning; case numbers match exactly.

  • Citations always. Every result and every answer points to the source file and page.

  • Always fresh. Folders are watched. Save a file and it is searchable in seconds.

  • Verifiably private. See below.

Trust commitments

These are versioned with the code and will not be quietly redrawn:

  1. The open/paid boundary is declared upfront. Open forever (AGPL-3.0): indexing, search, ask, REST API, MCP, CLI, and the upcoming single-user web UI. The full single-user product, no feature hostages. Paid (later): team features such as auth, roles, audit logs, shared indexes, and a hosted version.

  2. "Zero external connections" is a verifiable feature, not a claim. hay status reports outbound connections; the target is 0. No telemetry. If that ever changes it will be opt-in, documented loudly, and off by default.

  3. No silent network behavior. A local-first tool must never phone home to keep working.

  4. All development happens in this public repo. No private-repo surprises.

How it works

folders -> watcher -> extract (pdf/docx/pptx/md/txt/html/mbox) -> chunk -> embed
                                                            |
        you <- citations <- RRF merge <- FTS5 + vector search <- SQLite (one file)

Everything lives in a single SQLite database on your disk, and search is fully self-contained: the embedding model ships inside the binary. Answers (hay ask) are generated by whatever OpenAI-compatible local server you already run (Ollama, LM Studio, llama.cpp, Jan). Haypile itself ships no LLM and makes no network calls.

Set up a folder properly: hay init

For a folder you work in (a case folder, a project, a paper archive), hay init writes a per-folder config and wires everything up in one go:

cd ~/cases/acme-litigation
hay init          # three short questions, all with sensible defaults

It creates .haypile.yml (tag and exclude patterns), indexes the folder, optionally writes .mcp.json so Claude Code and Cursor can search these docs, and offers hay llm setup if you do not have a local LLM yet. hay init --yes runs unattended.

Edit .haypile.yml by hand anytime. The daemon notices and re-syncs the index within seconds:

tag: acme-litigation
exclude:
  - drafts/**
  - "*.bak"

Ask questions (bring your own LLM)

hay ask retrieves the most relevant passages and has a local LLM answer from them, with citations:

hay ask "what did the Meridian contract say about termination?"

Generation uses any OpenAI-compatible server you already run (Ollama, LM Studio, llama.cpp, Jan), auto-detected on their usual ports, or set explicitly with --endpoint and --model. Without one, hay ask explains and shows the top passages instead. Search never needs an LLM.

Prefer a cloud model for answers? Bring your own key:

hay ask --endpoint https://api.example.com/v1 --key sk-... "what changed in the lease?"

The boundary stays sharp: your documents are indexed locally, always. Opting in sends only the retrieved passages for that one question, to an endpoint you chose, with your key (HAYPILE_LLM_API_KEY works too). Keys are refused over plain http to anything that is not localhost.

No local LLM yet? One guided command gets you there:

hay llm setup    # installs and starts Ollama, pulls a model, asks before every download

Use from Claude Code, Cursor, or your own tools

The daemon exposes MCP (Streamable HTTP) and REST on localhost:11500:

# Claude Code
claude mcp add --transport http haypile http://localhost:11500/mcp

# Anything that prefers launching a process (stdio transport)
#   command: hay   args: ["mcp-stdio"]

# Plain REST
curl -X POST localhost:11500/api/query -d '{"query": "termination clause"}'

Tools exposed: search_documents (hybrid search with citations) and list_sources. The daemon starts automatically on hay add and only ever listens on localhost.

Commands

hay init [folder]        per-folder setup: config, index, editor wiring
hay add <path>           index a folder or file and watch it for changes
hay search "<query>"     hybrid retrieval, results with citations
hay ask "<question>"     answer from your documents, with cited sources
hay list                 indexed folders and document counts
hay remove <path>        un-index a folder
hay status               daemon state, model info, outbound connections (target: 0)
hay web                  open the local web UI in your browser
hay serve                run the daemon (REST API + MCP on localhost:11500)
hay llm setup            guided local LLM setup for hay ask

Roadmap

Version

Scope

v0.x (now)

CLI, REST API, MCP server, hay web local UI. Markdown, text, PDF, docx, pptx, HTML, mbox email. Scanned-PDF OCR via your local vision LLM.

v1.x

Bundled OCR (no LLM required), Windows installer polish

v2

Optional larger embedding models, ANN index for very large corpora

Pro

Team layer for offices: auth, roles, audit logs, shared indexes (paid)

Roadmap, decisions, and trust commitments in detail: docs/ROADMAP.md.

Development

go build ./cmd/hay     # build the binary
go test ./... -race    # run tests (green before any merge)

Semantic search uses an embedding model that release builds carry inside the binary. Dev builds load it from disk instead, so the weights stay out of git:

./hack/fetch-model.sh                    # one-time download (also quantizes)
go build -tags bundled ./cmd/hay         # release-style: model in the binary
HAYPILE_MODEL_PATH=internal/embed/bundled/model.safetensors ./hay   # dev

Without the model, everything still works in keyword-only mode.

The web UI (hay web) lives in webui/ as a small Vite + Preact app; its built output is committed under internal/webui/dist and embedded in the binary, so go build alone always ships the current UI. Touch the UI with:

cd webui && npm install
npm run dev      # live dev server, proxies /api to a running daemon
npm run build    # writes internal/webui/dist (commit the result)

Retrieval quality is measured, not vibes: eval/ holds a query set with expected results that runs on every retrieval-affecting change.

Contributing

Contributions are welcome. Please read CONTRIBUTING.md first. It covers the dev workflow, the rule that the privacy contract (zero outbound, bundled model, localhost only) must stay intact, how contributions are licensed, and the DCO sign-off (git commit -s) that CI enforces.

Security

Found a vulnerability? Please report it privately, not as a public issue. See SECURITY.md for the disclosure process and what is in scope.

License

AGPL-3.0. Free forever for individuals. The AGPL keeps it free for every actual user while requiring anyone offering Haypile as a service to open-source their changes.

Install Server
A
license - permissive license
A
quality
A
maintenance

Maintenance

Maintainers
Response time
1dRelease cycle
10Releases (12mo)
Commit activity

Related MCP Servers

  • A
    license
    -
    quality
    C
    maintenance
    Local offline semantic search over documents (txt, md, pdf, docx, pptx, csv). Indexes folders into a LanceDB vector database with multilingual embeddings and supports hybrid vector + keyword search via Reciprocal Rank Fusion. No API keys, no cloud, no Docker required.
    Last updated
    28
    AGPL 3.0
  • F
    license
    A
    quality
    D
    maintenance
    Enables indexing local documents (PDF, Markdown, text, code) into a knowledge base and querying them via semantic search using local embeddings, all running privately on your machine.
    Last updated
    4

View all related MCP servers

Related MCP Connectors

  • Search your knowledge bases from any AI assistant using hybrid RAG.

  • Local-first RAG engine with MCP server for AI agent integration.

  • Serve a folder of Markdown notes as an MCP server: hybrid search, reading, and sourced answers.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/BenyD/haypile'

If you have feedback or need assistance with the MCP directory API, please join our Discord server