Skip to main content
Glama
README.md
# memocean-mcp — Local-first AI Memory MCP Server

![MemOcean](assets/memocean-banner.jpg)

[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
[![Python 3.11+](https://img.shields.io/badge/Python-3.11%2B-blue.svg)](https://www.python.org/)

A local MCP server for AI agents to search and manage knowledge stored in Obsidian vaults.
Built for CJK (Chinese/Japanese/Korean) developers — 94.4% Hit@5 on Chinese queries with zero AI components required.

**Core features:**
- BM25/INSTR hybrid search — CJK-optimized, pure SQLite, no embeddings needed
- CLSC sonar compression — 92.5% token reduction (13x compression) on Obsidian notes
- Temporal knowledge graph — entity-relationship store with non-destructive invalidation
- Cross-agent memory sharing — multiple agents share one `memory.db`
- FATQ task queue — File-Atomic Task Queue for agent coordination

---

## Quick Start

```bash
pip install memocean-mcp
```

Add to Claude Desktop / Claude Code `.mcp.json`:

```json
{
  "mcpServers": {
    "memocean": {
      "command": "memocean-mcp",
      "env": {
        "MEMOCEAN_VAULT_ROOT": "/path/to/your/obsidian/vault"
      }
    }
  }
}
```

Or register with Claude Code CLI:

```bash
MEMOCEAN_VAULT_ROOT=/path/to/vault claude mcp add memocean memocean-mcp
```

---

## Environment Variables

| Variable | Default | Description |
|---|---|---|
| `MEMOCEAN_VAULT_ROOT` | `~/Documents/Obsidian Vault` | Root of your Obsidian vault |
| `MEMOCEAN_DATA_DIR` | `~/.memocean` | Data directory (databases, task queue) |
| `MEMOCEAN_VAULT_PATH` | `MEMOCEAN_VAULT_ROOT/Ocean` | Ocean subdirectory for full-text search |
| `MEMOCEAN_SKILLS_DIR` | `MEMOCEAN_VAULT_ROOT/Ocean/Pearl/skills` | Skills markdown directory |
| `MEMOCEAN_USE_GBRAIN` | `false` | Enable GBrain hybrid search delegate |
| `KNN_ENABLED` | `false` | Enable BGE-m3 KNN vector search |
| `ENABLE_QUERY_EXPANSION` | unset | Enable Haiku query expansion (requires `ANTHROPIC_API_KEY`) |
| `ENABLE_HAIKU_RERANKER` | unset | Enable Haiku LLM reranker |
| `ANTHROPIC_API_KEY` | unset | Required only for AI-assisted features above |

Backward-compat: `CHANNELLAB_BOTS_ROOT` → `MEMOCEAN_DATA_DIR`, `CHANNELLAB_OCEAN_VAULT_ROOT` → `MEMOCEAN_VAULT_ROOT`.

---

## Available Tools

| Tool | Description |
|---|---|
| `memocean_search` | Unified search across Radar (sonar index) + message history. Default entry point. |
| `memocean_radar_search` | Search CLSC sonar index — fast keyword search, ~13% of verbatim token cost. |
| `memocean_seabed_get` | Retrieve full content by slug (verbatim or sonar mode). |
| `memocean_ocean_search` | Full-text search over Ocean vault `.md` files via ripgrep. |
| `memocean_messages_search` | BM25 search over cross-agent message history. |
| `memocean_kg_query` | Query the temporal knowledge graph by entity name. |
| `memocean_skill_list` | List or retrieve approved skills from the skill library. |
| `memocean_task_create` | Create a task in the FATQ pending queue (agent coordination). |
| `memocean_ingest_file` | Ingest local file (PDF/DOCX/PPTX/XLSX/HTML/CSV/JSON) into Radar via MarkItDown. |
| `memocean_report_store` | Store a verbatim markdown report into Ocean vault Reports folder. |

---

## File-ingest dependency note

`memocean_ingest_file` depends on the PDF, DOCX, and PPTX extras from
MarkItDown. They are pinned in `pyproject.toml`; install the project dependencies
(for example, `pip install .`) rather than installing bare `markitdown`, which
omits the PDF and DOCX converters.

---

## Production deployment and drift guard

This host intentionally uses a normal package install, not an editable install.
Changing `memocean_mcp/` in this repository therefore does **not** change the
code imported from user site-packages. This separation prevents an uncommitted
worktree from silently becoming production behaviour.

After a reviewed memocean-mcp change is merged, an authorized production runner
must use the single operator-facing deployment entry:

```bash
MEMOCEAN_PIP_BREAK_SYSTEM_PACKAGES=1 \
  /home/oldrabbit/.claude-bots/shared/bin/memocean-deploy.sh
```

The deploy command refuses dirty package sources, performs a non-editable,
idempotent reinstall, runs the canonical content drift gate, and atomically
records the deployed repository commit in
`logs/memocean-mcp-deployment.json`. The deployment is successful only when
the final check reports `OK`. `ops/deploy.sh` remains the package-local
implementation used by the canonical entry and isolated fixture; it is not the
operator runbook entry.

Debian marks its base interpreter as externally managed (PEP 668), so pip also
requires the explicit `MEMOCEAN_PIP_BREAK_SYSTEM_PACKAGES=1` opt-in shown above.
The script applies pip's `--break-system-packages` only to the existing `--user`
install path; target-directory installs do not receive it. This does not write
into Debian's system package directory, but a user-site package can still
shadow a distro package for this account. Use this opt-in only on the reviewed
production host and interpreter. A dedicated virtual environment would avoid
that override, but adopting one also requires moving the service runtime and is
therefore a separate migration.

After review, install or refresh the recurring drift check with:

```bash
bash ops/install-drift-cron.sh
```

The installer creates one marked cron entry. The guard runs daily at 04:15
Asia/Taipei and appends a timestamped `trigger=cron` result to
`/home/oldrabbit/.claude-bots/logs/memocean-mcp-drift.log`. Immediately after
installation, run the check once with `MEMOCEAN_DRIFT_TRIGGER=cron` to verify
the cron-path output format, then retain the crontab entry and a fresh
scheduler-produced `trigger=cron` log line after the next scheduled run as
activation evidence. The manual run and crontab entry alone do not prove that
the scheduler is alive.

`ops/check-deploy-drift.sh` is a compatibility shim for the already deployed
daily cron; it delegates to `shared/bin/memocean-drift-check.sh`, the same
canonical gate used by deploy and the FATQ live probe. When drift exists, the
guard also uses the existing `mm_post` path to notify
Anya's Mattermost channel. An unchanged drift signature is sent once across
consecutive daily runs; a changed signature is sent again, and a clean run
resets the deduplication state. Notification failure is printed explicitly and
never changes the drift result from exit 1. Host activation evidence must
include an injected-drift message that was actually received, not only script
configuration or a log line.

The drift guard compares only `.py` source files. It deliberately ignores
`__pycache__`, `.pyc`, build metadata, and generated bytecode. A missing, extra,
or changed `.py` file is listed and makes the command exit nonzero.

---

## Search Architecture

Two-path retrieval, zero AI dependency by default:

```
CJK query  →  SQLite INSTR on radar.clsc  →  ranked by match_count
EN query   →  FTS5 BM25                   →  fallback to INSTR on miss
```

Benchmark (pure BM25/INSTR, no AI):

| Dataset | Language | Hit@5 |
|---|---|---|
| Internal corpus | Chinese (mixed) | **94.4%** |
| DRCD | Traditional Chinese | **91.9%** |
| CMRC | Simplified Chinese | **93.3%** |
| BEIR SciFact | English | 70.7% |

---

## CLSC Sonar Compression

CLSC (Closet Lossy Summary for Chinese) extracts each document into a compact single-line sonar entry. Format:

```
[SLUG|ENTITIES|topics|"key_quote"|WEIGHT|EMOTIONS|FLAGS]
```

Compression ratio: **1,716,211 raw tokens → 129,529 sonar tokens = 13x (92.5% reduction)**.

---

## Requirements

- Python 3.11+
- SQLite 3.35+
- Optional: `markitdown[all]` for file ingestion
- Optional: `anthropic` package for AI-assisted features (query expansion, reranking)

---

## License

MIT — see [LICENSE](LICENSE).

---

## Acknowledgements

Built on [MemPalace](https://github.com/milla-jovovich/mempalace) (dual-layer architecture, AAAK skeleton format) and inspired by [GBrain](https://github.com/garrytan/gbrain) (Compiled Truth + Dream Cycle design).