Lexomni MCP
by PagansDev
README.md
# Lexomni MCP
MCP (Model Context Protocol) server for local knowledge management with Markdown and PDF indexing using SQLite FTS5.
**Author:** [PagansDev](https://github.com/PagansDev) (Paulo Gabriel Neves Santos) | **Repository:** [github.com/PagansDev/lexomni-mcp](https://github.com/PagansDev/lexomni-mcp)
---
## English
### Usage
Requires **Node.js 20.16 or newer**.
#### As a local MCP server
1. Configure your MCP client (e.g., `.mcp.json` for Claude Code, `.cursor/mcp.json` for Cursor). Pin an exact version:
```json
"lexomni": {
"command": "npx",
"args": ["-y", "lexomni-mcp@0.2.0"],
"env": { "LEXOMNI_WORKSPACE": "_lexomni" }
}
```
A relative `LEXOMNI_WORKSPACE` keeps a versioned config file portable across machines. Projects set up with [`@satlab/protocol`](https://www.npmjs.com/package/@satlab/protocol) get this entry from `satlab-protocol init`, and `satlab-protocol update` keeps it current.
2. Add your documents:
- `_lexomni/user/` - User markdown (guidelines, architecture, etc.)
- `_lexomni/agent/` - Agent notes (written by `lexomni_writeNote`)
- `_lexomni/books/` - PDFs (books, documentation, etc.)
3. Ask the agent to run `lexomni_buildIndex`. Searches only return what has been indexed.
#### How the workspace is resolved
1. `LEXOMNI_WORKSPACE`, when set. An absolute path is used as is; a relative one is resolved against the project root.
2. The project root is `CLAUDE_PROJECT_DIR` (set by Claude Code), otherwise the first root reported by the MCP client, otherwise the current directory.
3. Without `LEXOMNI_WORKSPACE`, the server looks for a `_lexomni` folder from the project root upwards and **stops at the repository root** (the first folder containing `.git`). If none is found, the workspace is `<project root>/_lexomni`.
Reading never creates folders: `user/`, `agent/`, `books/` and `index/` are created only when `lexomni_writeNote` or `lexomni_buildIndex` writes. `lexomni_listSources` and `lexomni_searchDocs` always report the workspace they used and return a `warning` when it is empty, when the index does not exist, or when it is out of date.
### Available Tools
> Agent handles parameters on its own
#### `lexomni_listSources`
Lists all documents (MD and PDF) in the workspace.
**Parameters:** none
**Example:**
```json
{
"workspace": "/path/to/project/_lexomni",
"count": 5,
"docs": [...]
}
```
#### `lexomni_buildIndex`
Indexes documents into SQLite FTS5 for fast search. This is the **only** tool that indexes, and it is incremental: documents whose size and modification time did not change are skipped, and documents removed from the workspace are dropped from the index.
**Parameters:**
- `sources` (optional): array of `["user", "agent", "books"]`. Limits what is indexed and what can be removed.
- `force` (optional): reindex everything, changed or not.
**Example:**
```json
{ "ok": true, "docsIndexed": 1, "indexed": 1, "skipped": 11, "removed": 0 }
```
#### `lexomni_searchDocs`
Keyword search across indexed documents. It never indexes: new or changed files appear only after the next `lexomni_buildIndex`, and the response carries a `warning` until then.
The query is literal text. Each word is matched as typed, so hyphens, quotes and colons are safe (`third-party`, `CVE-2024-4367`); FTS5 operators such as `OR`, `NEAR` or `*` are not interpreted.
**Parameters:**
- `query` (required): string, min 2 characters
- `sources` (optional): filter by source
- `limit` (optional): max results (1-50, default 10)
**Example:**
```json
{
"query": "clean architecture",
"workspace": "/path/to/project/_lexomni",
"hits": [
{
"docId": "user:user/guidelines.md",
"chunkIndex": 0,
"snippet": "...about [clean] [architecture]...",
"source": "user",
"relPath": "user/guidelines.md"
}
]
}
```
#### `lexomni_readDoc`
- docId (required)
- Type: string
Purpose: unique identifier of the document in the index.
> Constraint: at least 3 characters.
Typical source: taken from a hit returned by lexomni_searchDocs.
- chunkIndex (optional)
Type: integer
Minimum: 0
Purpose: which chunk (piece) of the document to read.
If omitted: usually defaults to the first chunk (0), depending on implementation.
- maxChars (optional)
Type: integer
Range: 200–20000
Purpose: maximum number of characters of text to return for that chunk, useful to limit response size.
#### `lexomni_writeNote`
- filename (required)
Type: string
Purpose: name of the markdown file to create or update in the agent’s notes area.
> Constraint: at least 1 character.
- content (required)
Type: string
Purpose: markdown content to write into the file.
> Constraint: at least 1 character.
- mode (optional)
Type: string
Allowed values:
"overwrite" – replaces the existing file content entirely.
"append" – appends content to the end of the existing file.
Default: "overwrite" if not specified.
### Security
- Path traversal blocked (no `../`, absolute paths or null bytes in note names)
- Write access limited to `_lexomni/agent/`, checked against the real path on disk
- Symbolic links are ignored when indexing and refused when writing notes, so a symlink inside the workspace cannot expose or overwrite files outside it
- PDFs are parsed only by `lexomni_buildIndex`, each one in a separate, sandboxed Node process (read-only access to the package and that PDF, no writes, no child processes, 100 MB / 512 MB / 60 s limits), with pdf.js string evaluation disabled
- Notes are limited to 1 MiB
- Runtime dependencies are bundled and releases are published from CI with provenance
See [SECURITY.md](SECURITY.md) for the supply chain policy, risky components and how to report a vulnerability.
### Architecture
```
_lexomni/ # In user's project
user/ # Guidelines, architecture (read-only)
agent/ # Agent notes (writable)
books/ # PDFs (read-only)
index/ # SQLite FTS5 (generated)
lexomni.sqlite
```
### Multilingual Strategy
Lexomni does not translate text internally. To find docs in PT and EN for e.g. :
1. Agent expands queries before searching:
- "arquitetura em camadas" → also searches "layered architecture"
- "fila" → also searches "message queue", "job queue"
2. Living glossary at `_lexomni/agent/glossary.md`:
- Agent learns new terms via web search
- Improves search quality over time
### Development
```bash
npm run build # Bundle into dist/ (index.js, pdf.worker.mjs, sql-wasm.wasm)
npm run dev # Watch mode
npm start # Run built server
npm test # Tests
npm run lint # Lint
npm run typecheck # Typecheck
```
To run a local build from a clone, point the MCP client to it:
```json
"lexomni-local": {
"command": "node",
"args": ["/path/to/lexomni-mcp/dist/index.js"],
"env": { "LEXOMNI_WORKSPACE": "_lexomni" }
}
```
---
## License
MIT © [Paulo Gabriel Neves Santos](https://github.com/PagansDev)
TDQS
A3.7/5.0
Scored across 5 tools
Disambiguation5/5
Each tool has a clearly distinct function: listing sources, building the index, searching, reading a chunk, and writing a note. There is no overlap or confusion between tool purposes.
Naming Consistency5/5
All tools follow a consistent lexomni_verbNoun pattern (listSources, buildIndex, searchDocs, readDoc, writeNote). The naming is uniform and predictable.
Tool Count5/5
Five tools is well-scoped for a document indexing and retrieval server. Each tool earns its place, covering the essential operations without unnecessary bloat.
Completeness4/5
Core workflows are covered: list sources, build index, search, read chunks, and write agent notes. A minor gap is the lack of a way to enumerate chunks or retrieve document structure, but this is workable for the intended use case.
Maintenance
ActivityMaintained
ResponsivenessUnresponsive