Skip to main content
Glama
README.md
# Lexomni MCP

MCP (Model Context Protocol) server for local knowledge management with Markdown and PDF indexing using SQLite FTS5.

**Author:** [PagansDev](https://github.com/PagansDev) (Paulo Gabriel Neves Santos) | **Repository:** [github.com/PagansDev/lexomni-mcp](https://github.com/PagansDev/lexomni-mcp)

---

## English

### Usage

Requires **Node.js 20.16 or newer**.

#### As a local MCP server

1. Configure your MCP client (e.g., `.mcp.json` for Claude Code, `.cursor/mcp.json` for Cursor). Pin an exact version:

```json
"lexomni": {
  "command": "npx",
  "args": ["-y", "lexomni-mcp@0.2.0"],
  "env": { "LEXOMNI_WORKSPACE": "_lexomni" }
}
```

A relative `LEXOMNI_WORKSPACE` keeps a versioned config file portable across machines. Projects set up with [`@satlab/protocol`](https://www.npmjs.com/package/@satlab/protocol) get this entry from `satlab-protocol init`, and `satlab-protocol update` keeps it current.

2. Add your documents:
   - `_lexomni/user/` - User markdown (guidelines, architecture, etc.)
   - `_lexomni/agent/` - Agent notes (written by `lexomni_writeNote`)
   - `_lexomni/books/` - PDFs (books, documentation, etc.)

3. Ask the agent to run `lexomni_buildIndex`. Searches only return what has been indexed.

#### How the workspace is resolved

1. `LEXOMNI_WORKSPACE`, when set. An absolute path is used as is; a relative one is resolved against the project root.
2. The project root is `CLAUDE_PROJECT_DIR` (set by Claude Code), otherwise the first root reported by the MCP client, otherwise the current directory.
3. Without `LEXOMNI_WORKSPACE`, the server looks for a `_lexomni` folder from the project root upwards and **stops at the repository root** (the first folder containing `.git`). If none is found, the workspace is `<project root>/_lexomni`.

Reading never creates folders: `user/`, `agent/`, `books/` and `index/` are created only when `lexomni_writeNote` or `lexomni_buildIndex` writes. `lexomni_listSources` and `lexomni_searchDocs` always report the workspace they used and return a `warning` when it is empty, when the index does not exist, or when it is out of date.

### Available Tools
> Agent handles parameters on its own

#### `lexomni_listSources`
Lists all documents (MD and PDF) in the workspace.

**Parameters:** none

**Example:**
```json
{
  "workspace": "/path/to/project/_lexomni",
  "count": 5,
  "docs": [...]
}
```

#### `lexomni_buildIndex`
Indexes documents into SQLite FTS5 for fast search. This is the **only** tool that indexes, and it is incremental: documents whose size and modification time did not change are skipped, and documents removed from the workspace are dropped from the index.

**Parameters:**
- `sources` (optional): array of `["user", "agent", "books"]`. Limits what is indexed and what can be removed.
- `force` (optional): reindex everything, changed or not.

**Example:**
```json
{ "ok": true, "docsIndexed": 1, "indexed": 1, "skipped": 11, "removed": 0 }
```

#### `lexomni_searchDocs`
Keyword search across indexed documents. It never indexes: new or changed files appear only after the next `lexomni_buildIndex`, and the response carries a `warning` until then.

The query is literal text. Each word is matched as typed, so hyphens, quotes and colons are safe (`third-party`, `CVE-2024-4367`); FTS5 operators such as `OR`, `NEAR` or `*` are not interpreted.

**Parameters:**
- `query` (required): string, min 2 characters
- `sources` (optional): filter by source
- `limit` (optional): max results (1-50, default 10)

**Example:**
```json
{
  "query": "clean architecture",
  "workspace": "/path/to/project/_lexomni",
  "hits": [
    {
      "docId": "user:user/guidelines.md",
      "chunkIndex": 0,
      "snippet": "...about [clean] [architecture]...",
      "source": "user",
      "relPath": "user/guidelines.md"
    }
  ]
}
```
#### `lexomni_readDoc`
- docId (required)
- Type: string
  Purpose: unique identifier of the document in the index.
> Constraint: at least 3 characters.
  Typical source: taken from a hit returned by lexomni_searchDocs.
- chunkIndex (optional)
  Type: integer
  Minimum: 0
  Purpose: which chunk (piece) of the document to read.
  If omitted: usually defaults to the first chunk (0), depending on implementation.
- maxChars (optional)
  Type: integer
  Range: 200–20000
  Purpose: maximum number of characters of text to return for that chunk, useful to limit response size.
  
#### `lexomni_writeNote`
- filename (required)
  Type: string
  Purpose: name of the markdown file to create or update in the agent’s notes area.
> Constraint: at least 1 character.
- content (required)
  Type: string
  Purpose: markdown content to write into the file.
> Constraint: at least 1 character.
- mode (optional)
  Type: string
  Allowed values:
  "overwrite" – replaces the existing file content entirely.
  "append" – appends content to the end of the existing file.
  Default: "overwrite" if not specified.

### Security

- Path traversal blocked (no `../`, absolute paths or null bytes in note names)
- Write access limited to `_lexomni/agent/`, checked against the real path on disk
- Symbolic links are ignored when indexing and refused when writing notes, so a symlink inside the workspace cannot expose or overwrite files outside it
- PDFs are parsed only by `lexomni_buildIndex`, each one in a separate, sandboxed Node process (read-only access to the package and that PDF, no writes, no child processes, 100 MB / 512 MB / 60 s limits), with pdf.js string evaluation disabled
- Notes are limited to 1 MiB
- Runtime dependencies are bundled and releases are published from CI with provenance

See [SECURITY.md](SECURITY.md) for the supply chain policy, risky components and how to report a vulnerability.

### Architecture

```
_lexomni/               # In user's project
  user/                 # Guidelines, architecture (read-only)
  agent/                # Agent notes (writable)
  books/                # PDFs (read-only)
  index/                # SQLite FTS5 (generated)
    lexomni.sqlite
```

### Multilingual Strategy

Lexomni does not translate text internally. To find docs in PT and EN for e.g. :

1. Agent expands queries before searching:
   - "arquitetura em camadas" → also searches "layered architecture"
   - "fila" → also searches "message queue", "job queue"

2. Living glossary at `_lexomni/agent/glossary.md`:
   - Agent learns new terms via web search
   - Improves search quality over time

### Development

```bash
npm run build      # Bundle into dist/ (index.js, pdf.worker.mjs, sql-wasm.wasm)
npm run dev        # Watch mode
npm start          # Run built server
npm test           # Tests
npm run lint       # Lint
npm run typecheck  # Typecheck
```

To run a local build from a clone, point the MCP client to it:

```json
"lexomni-local": {
  "command": "node",
  "args": ["/path/to/lexomni-mcp/dist/index.js"],
  "env": { "LEXOMNI_WORKSPACE": "_lexomni" }
}
```

---

## License

MIT © [Paulo Gabriel Neves Santos](https://github.com/PagansDev)

TDQS

A3.7/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct function: listing sources, building the index, searching, reading a chunk, and writing a note. There is no overlap or confusion between tool purposes.

Naming Consistency5/5

All tools follow a consistent lexomni_verbNoun pattern (listSources, buildIndex, searchDocs, readDoc, writeNote). The naming is uniform and predictable.

Tool Count5/5

Five tools is well-scoped for a document indexing and retrieval server. Each tool earns its place, covering the essential operations without unnecessary bloat.

Completeness4/5

Core workflows are covered: list sources, build index, search, read chunks, and write agent notes. A minor gap is the lack of a way to enumerate chunks or retrieve document structure, but this is workable for the intended use case.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive