Skip to main content
Glama
noualit

llama-memory

by noualit

llama-memory

llama-server 用の MCP メモリサービス。Postgres + PGVector による永続的な履歴とセマンティックメモリを提供します。

注: このプロジェクトはローカル/デモ用途のみを想定しています。追加の堅牢化(HTTPS、適切な認証、バックアップ)なしでインターネットに直接公開しないでください。

機能

  • セマンティックメモリ: キーワードだけでなく意味に基づいてメモリを保存・取得します。

  • 会話ブリッジ: LLM が会話を自動的に作成し、メモリがリンクされます。

  • セッションをまたいだ想起: 「前に何を話しましたか?」と尋ねると正確な回答が得られます。

  • MCP プロトコル: llama-server の組み込み MCP サポートと直接連携します。

Related MCP server: engram

要件

  • Python 3.11(Miniconda 推奨)

  • PGVector 拡張機能を備えた PostgreSQL 16+

  • --jinja フラグ付きの llama-server(ツール呼び出しに必要)

  • llama-server 上で動作する nomic-embed-text(デフォルトではポート 8081)

インストール

# Clone the repo
git clone https://github.com/noualit/llama-memory-local.git
cd llama-memory-local

# Create environment
conda create -n llama-memory python=3.11
conda activate llama-memory

# Install dependencies
pip install -e .

設定

.env.example を .env にコピーして編集します:

cp .env.example .env

例:

# Database
DATABASE_URL="postgresql://postgres:yourpassword@localhost:5432/llamamem"

# Llama-server (LLM)
LLAMA_SERVER_BASE_URL="http://localhost:8080"

# Embedding model (nomic-embed-text via llama-server)
EMBEDDING_MODEL_URL="http://localhost:8081"

# Embedding model name (default: nomic-embed-text)
EMBEDDING_MODEL_NAME="nomic-embed-text"

# Service port
SERVICE_PORT=9001

データベースのセットアップ

データベースを作成し、マイグレーションを実行します:

psql -U postgres -c "CREATE DATABASE llamamem;"
alembic upgrade head

アプリケーションは起動時に基本スキーマも自動的に確保します。

サービスの実行

# Using the script
.\scripts\run_server.ps1

# Or directly
python -m uvicorn app.main:app --host 0.0.0.0 --port 9001

llama-server への接続

llama-server の MCP 設定に追加します:

{
  "mcpServers": {
    "llama-memory": {
      "url": "http://YOUR_SERVER_IP:9001/mcp"
    }
  }
}

サービスは llama-server から到達可能である必要があります。別々のマシンで実行している場合は、localhost ではなく実際の IP を使用してください。

MCP ツール

ツール

説明

create_conversation

新しい会話セッションを作成

list_conversations

メモリ数付きで会話を一覧表示

get_conversation_history

会話内のすべてのメモリを取得

search_memories

すべてのメモリをセマンティック検索

save_memory

重要な事実や決定を保存

システムプロンプト

次のことができます:

  • サービスから推奨システムプロンプトを取得:

    • GET /system-prompt → プレーンテキストを返します。

  • または、この最小バージョンを llama-server に貼り付けます:

MEMORY WORKFLOW:
- At the start of each new conversation, call create_conversation with a short title.
- Use the conversation_id from create_conversation when calling save_memory.
- Before answering questions about past topics, call search_memories FIRST.
- When the user shares important information, save it with save_memory.
- If list_conversations has previous chats, check get_conversation_history for context.

ヘルスチェック

curl http://localhost:9001/health

DB のステータス、埋め込みサービスのステータス、ツール数を返します。

アーキテクチャ

高レベルの構造:

  • app/main.py — FastAPI アプリ、ライフスパン、/system-prompt

  • app/settings.py — .env からの Pydantic 設定

  • app/clients/embeddings.py — ベクトル用に nomic-embed-text を呼び出し

  • app/db/engine.py — asyncpg 接続プール(シングルトン)

  • app/db/schema.py — 起動時にテーブルを自動作成

  • app/mcp/endpoint.py — MCP プロトコルハンドラー、レートリミッター

  • app/mcp/tools/ — 個々のツール実装

  • migrations/ — Alembic データベースマイグレーション

開発

# Run tests
pytest tests/ -v

# Run with auto-reload
python -m uvicorn app.main:app --host 0.0.0.0 --port 9001 --reload

貢献ガイドラインについては CONTRIBUTING.md を参照してください。

ライセンス

MIT(LICENSE ファイルを参照)。

A
license - permissive license
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Persistent semantic memory server for AI assistants via MCP, enabling long-term context retention and semantic search across conversations.
    11
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides persistent, local-first AI memory across sessions via MCP tools for storing, searching, and retrieving context from past interactions.
    1
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Provides persistent memory for AI assistants via MCP, enabling them to store and recall facts, preferences, and tasks across conversations using either local file storage or a cloud backend with semantic search.
    5
    14
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Provides persistent memory with semantic search for MCP-based AI agents, enabling them to store and recall information across sessions using vector embeddings.
    4
    1
    MIT

View all related MCP servers

Related MCP Connectors

  • Shared, governed long-term memory for AI agents across tools and sessions via MCP and REST.

  • Persistent memory for AI agents. Search, store, and recall across sessions.

  • Cross-AI personal memory. Save once in ChatGPT, recall in Claude, Mistral, Grok, or any MCP client.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/noualit/llama-memory-local'

If you have feedback or need assistance with the MCP directory API, please join our Discord server