Skip to main content
Glama
hermes-labs-ai

io.github.hermes-labs-ai/fidelis-memory

Official

Fidelis Memory

Agent memory that brings back the source, not another summary.

Fidelis is developed by Hermes Labs.

Hermes Labs is an agentic infrastructure company building the reliability layer for autonomous systems.

PyPI pre-release CI Python License: Apache-2.0

Fidelis is a local memory and retrieval service for Codex, Claude Code, and other AI agents. Keep your notes available across sessions and retrieve stored text without generative rewriting.

A summary can preserve "we tried the migration" while dropping why it failed, what it affected, and what must change before trying again. Fidelis's verbatim ingestion path keeps those details in the stored note instead of requiring a generated fact to replace it.

Quickstart · Connect your agent · How it works · Benchmarks · Documentation

Quickstart

You need Python 3.10+, macOS or Ubuntu, and Ollama running locally. Ubuntu service installation uses systemd. Install Ollama first; if its server is not running, start ollama serve in another terminal. This walkthrough needs no model API key.

1. Install and start Fidelis

ollama pull nomic-embed-text

python3 -m venv ~/.venvs/fidelis
source ~/.venvs/fidelis/bin/activate
python3 -m pip install "fidelis-memory[hybrid]==0.3.0rc1"

fidelis init

The hybrid extra adds BM25 keyword search. fidelis init installs the background memory service using this Python environment, so keep the environment in place. The package is fidelis-memory; the command is fidelis.

2. Store a note and retrieve it

demo_dir=$(mktemp -d)
cat > "$demo_dir/atlas.md" <<'NOTE'
Atlas billing migration, 2026-09-20:
Duplicate charges appeared in staging. Rolled back.
Do not retry until the idempotency fix is verified.
NOTE

fidelis watch "$demo_dir" --once
fidelis recall-hybrid "Atlas billing migration retry condition" --tier zero_llm

Success means the returned text includes both the rollback and the retry condition. The command retrieves stored text; it does not generate an answer. Scores and ordering depend on your store.

For your own notes, run fidelis watch ~/notes --once. Omit --once to keep watching in a separate terminal. The watcher ingests Markdown and text files, not every conversation in your agent clients.

Trouble retrieving? Run fidelis health and check that Ollama is running with nomic-embed-text available. A responding health endpoint alone does not prove that ingestion and retrieval work.

Related MCP server: cc-history

Connect your agent

After the local retrieval works, register Fidelis with the client you use:

Client

Install command

Remove command

Codex

fidelis mcp install --client codex

fidelis mcp uninstall --client codex

Claude Code

fidelis mcp install

fidelis mcp uninstall

Cursor

fidelis mcp install --client cursor

fidelis mcp uninstall --client cursor

GitHub Copilot CLI

fidelis mcp install --client copilot

fidelis mcp uninstall --client copilot

Gemini CLI

fidelis mcp install --client gemini

fidelis mcp uninstall --client gemini

OpenClaw

fidelis mcp install --client openclaw

fidelis mcp uninstall --client openclaw

The installer preserves other MCP servers and refuses to replace a different server named fidelis unless you explicitly use --force. Cursor's default destination is ~/.cursor/mcp.json; pass --settings PATH to target a project .cursor/mcp.json instead. Keep the Python environment used to install Fidelis in place, because Cursor launches that environment's bundled MCP server.

Restart your client, confirm fidelis appears in its MCP tool list (Cursor: Customize → MCP), then try:

Use Fidelis to retrieve my Atlas billing migration note. What must happen before we retry? Quote the relevant text.

Fidelis exposes six MCP tools: fidelis_recall, fidelis_store, fidelis_correct, fidelis_get, fidelis_recent, and fidelis_health. Ask for fidelis_health first: it distinguishes a registered client from a reachable local service. Then fidelis_recall should return the Atlas note, including the exact retry condition. Your agent decides when to call the tools; registration does not guarantee automatic recall on every turn. See the technical reference for client prerequisites and configuration.

The repository root also supplies a portable Agent Plugin (plugin.json and mcp.json) for clients that load Agent Plugins 1.0, including Cursor. It launches the same released stdio MCP server through uvx; install uv first. Use either that plugin or fidelis mcp install --client cursor in one Cursor profile, to avoid two copies of the six tools. The local service and ingested notes are still required. The plugin adds no persistent memory by itself.

Pi prompt-time recall

Pi 0.87.1 or newer (Node.js 22.19 or newer) can load the repository's native extension after the local Fidelis service and your notes are ready:

pi install git:github.com/hermes-labs-ai/fidelis@main
pi list

Restart Pi or run /reload. Installing this Git package opts in to one local POST /recall_b before each user turn containing at least three non-whitespace characters. The extension sends only the expanded prompt to 127.0.0.1 on FIDELIS_PORT (or COGITO_PORT, default 19420), requests at most three results, and displays their source text in the Pi transcript before the model answers. Long notes appear as marked, exact prefixes with their IDs so you can retrieve the full record. Older recall messages remain visible in the transcript but leave the next turn's model context; all recall messages are excluded from compaction summaries. The extension never writes memory. Unavailable or slow recall produces a warning and lets the turn continue within one second. An empty result adds no context. Pi package registration alone does not prove recall worked: ask about a distinctive note you have already ingested and confirm its exact text appears in the displayed fidelis-pi-recall message.

This route adds prompt-time context, not the six MCP tools or a new memory store. You can use Pi's MCP adapter separately when you need explicit get, store, or correction tools. /recall_b does not apply Fidelis's full temporal view, so this adapter labels temporal status as unchecked. Retrieved notes may be outdated or superseded; inspect their status and source before relying on them. Remove this adapter with pi remove git:github.com/hermes-labs-ai/fidelis@main.

MCP update in 0.3.0rc1: recall, recent results, and correction chains return full stored text; the old silent 300-character previews are removed. Corrections retain superseded records, and recall supports validity dates and historical views. Replace old fidelis_query calls with fidelis_recall and restart clients to refresh their tool lists. See the upgrade and rollback notes.

Why keep the source?

Summaries are useful for navigating a long history. They can also leave out information that becomes important to a later question. Once the summary is all that remains, retrieval cannot recover what was discarded.

Fidelis is built for work where you need to revisit the evidence:

  • Decisions and constraints: recover the rationale, exceptions, and exact conditions in a saved note.

  • Failed approaches: retrieve what broke and what must change before another attempt.

  • Work across sessions: make your saved project context accessible to different agent clients on the same machine.

The principle is simple: use derived representations to find evidence, not to replace it.

How it works

Your Markdown or text files
          |
   Verbatim ingestion
          |
   Local memory store
          |
   Retrieve and rank candidates
          |
   Stored text for your agent

Default MCP recall uses fast local vector retrieval without a generative LLM. Explicit mode: "thorough" selects the hybrid path. The hybrid retrieval path combines keyword search, dense-vector similarity, and reciprocal rank fusion. Its default zero_llm tier does not call a generative LLM. Local embeddings are still required.

Optional model-assisted tiers can help select candidates. Their accepted output is a list of candidate numbers. Code resolves those numbers to stored text rather than returning the model's prose as memory.

Fidelis builds on mem0 and ChromaDB for storage and adds its retrieval, fidelity, service, and agent-integration layers.

The fidelity boundary

The source-preserving paths include fidelis watch, fidelis store, fidelis add, and HTTP POST /store. Explicit fidelis add --extract and fidelis seed use extraction or curation and can transform input before storage. Snapshots are derived summaries, not source evidence.

Fidelity means preserving the text supplied through the verbatim path. It does not prove that the text is true, current, complete, or the original record of an event. Store only a summary and only that summary can be recovered. Full provenance tracking is not a release guarantee.

With local Ollama, the quickstart keeps storage and retrieval local. Your agent may send retrieved text to its model provider when answering. Optional LLM features follow their configured data boundaries.

Benchmarks

The redesigned default zero-LLM retrieval path completed a fresh LongMemEval-S run of 470 questions on September 21, 2026, with zero errors. The results and methodology record the source snapshot, corpus construction, metrics, and limitations.

These are whole-session retrieval measurements, not answer accuracy or a matched comparison with competitors. Historical chunked retrieval and QA scores do not measure this redesign and are not reused as release evidence. Full LLM/QA evaluation is post-release work.

Fast recall remains the default. Optional thorough hybrid retrieval needs further tuning; that work is deferred beyond this pre-release.

Is Fidelis a fit?

Choose Fidelis when your working context lives in local notes, you want to retrieve their text rather than replace it with synthesized memory, and you can run a local service.

Version 0.3.0rc1 is an early, single-machine pre-release. It is not a hosted team-memory platform. Windows service installation and managed multi-user authorization are not supported contracts. Preserving a past statement also does not make it current: review dates and conflicting records before acting.

See the user-fit guide and security policy before deploying.

Documentation

Need

Start here

Commands, HTTP API, configuration, and client setup

Technical reference

Supported workflows and limitations

User-fit guide

Optional guidance for an LLM reading retrieved evidence

QA scaffold

Release history

Changelog

Contributing or reporting security issues

Contributing · Security

Contributing

Found a missed passage, an unexpected rewrite, or an installation problem? Open an issue with a minimal, redacted example. Retrieval regressions, fidelity tests, and documentation fixes are welcome.

License

Apache-2.0.

Available Tools

4 tools
fidelis_healthB

Check the fidelis server's health and memory count.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral burden. It accurately implies a read-only health-check operation, but it does not disclose output format, what 'memory count' refers to, or any side effects or permissions. This is adequate but minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It states the action and object directly and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool, invocation is trivial, so no parameter guidance is missing. However, with no output schema and no annotations, the description should clarify what 'health and memory count' actually return, such as status values, units, or response shape. The ambiguity leaves the agent without a complete picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is empty, so the no-parameter baseline of 4 applies. The description adds no parameter-level detail, but none is needed here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Check') and names the resource ('fidelis server's health and memory count'), making the tool's role clear. It distinguishes itself from the sibling tools, which sound like data operations (recall, query, orient), though it does not explicitly contrast itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus the sibling tools, nor are there exclusions, prerequisites, or context cues. The usage is only implied by the name and brief description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fidelis_orientA

Context-sensitive re-entry for prior work. Call when a turn mentions a known project, decision, earlier work, maintenance, comparison, or possible reuse—even when the turn is not phrased as a question. Returns an evidence-bound orientation packet or explicitly abstains.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
entityNoOptional exact known project or concept name
utteranceYesCurrent user turn
recent_turnsNoUp to four recent turns for referent resolution

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does disclose a meaningful behavioral trait: it returns an evidence-bound orientation packet or explicitly abstains. However, it does not clarify what the packet contains, whether the operation is read-only, or what triggers abstention beyond the listed topics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences with no filler. The purpose, call conditions, and output/abstention behavior are compactly and effectively ordered; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without annotations or an output schema, the description gives a clear high-level contract but leaves the orientation packet contents and the role of optional parameters under-specified. It is enough to guide an agent toward the tool, but not enough to fully set expectations about what comes back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already documents utterance, entity, and recent_turns. The description adds usage context for the utterance but does not clarify the limit parameter or define the evidence-bound packet contents. It neither contradicts nor meaningfully augments the schema, leaving the schema to carry most of the weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific function: context-sensitive re-entry into prior work, and lists concrete trigger signals (known project, decision, earlier work, maintenance, comparison, reuse). It does not explicitly distinguish itself from sibling tools fidelis_recall, fidelis_query, or fidelis_health, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-call guidance, covering a wide range of trigger phrases and noting that it applies even when the turn is not phrased as a question. It does not name alternatives or state when not to use this tool, so it lacks the explicit exclusion that would make a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fidelis_queryB

Fast vector-only query over fidelis memories (no filter).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results
queryYesSearch query

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It mentions 'fast' and 'vector-only', which conveys performance and method, but it does not state that the operation is read-only or non-destructive, nor any other side effects, authentication needs, or limitations. For a query tool this is a notable omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler words. It front-loads the core purpose and includes the most critical limitation ('no filter'). Every word contributes meaning, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter tool with no output schema, the description covers the essential action and scope. However, given the absence of annotations, it fails to mention that the operation is safe/read-only, and it does not describe the return format. These gaps make it less complete than it could be.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — both 'limit' and 'query' have basic descriptions. The description adds context about the overall query behavior ('vector-only', 'no filter') but does not provide additional per-parameter meaning beyond the schema. Baseline 3 applies because the schema already documents the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a verb (query) and a resource (fidelis memories), and adds specific qualifiers ('vector-only', 'no filter') that distinguish it from a filtered query. It does not explicitly name sibling tools, but the qualifiers provide enough differentiation for an agent to guess it is a basic, unfiltered search variant.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when no filter is needed ('no filter'), which gives a condition, but it does not explicitly mention when to prefer siblings like fidelis_recall or fidelis_orient. The guidance is implicit rather than explicit, so an agent must infer that filtered or more advanced queries belong elsewhere.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fidelis_recallB

Retrieve memories from the local fidelis store. Two-stage recall with optional LLM filter; zero-LLM by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results
queryYesNatural-language query

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral disclosure burden. It does add useful context — 'Two-stage recall with optional LLM filter; zero-LLM by default' — revealing internal mechanics beyond a simple retrieval. However, it omits any clarification of the two stages, possible side effects, or output shape, and the optional LLM filter is not backed by any parameter in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that states the action and key behavioral constraints with zero waste. Every clause adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema and no annotations, so the description must carry more weight. It fails to explain how the optional LLM filter is toggled (no schema parameter exists for it), what 'two-stage' entails, or how this differs from fidelis_query. This leaves an agent with a meaningful gap in knowing how to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'query' and 'limit' already documented meaningfully. The description adds no additional parameter-level meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Retrieve') and resource ('memories from the local fidelis store'), clearly indicating a read operation. However, it does not explicitly differentiate itself from the sibling tool 'fidelis_query', so an agent may struggle to choose between them without further inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus fidelis_query or fidelis_orient. It mentions 'zero-LLM by default' as a behavioral trait, but does not state conditions such as 'use when you need deterministic recall' or 'avoid when...' This leaves selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedfidelis_health
    • First observedfidelis_orient
    • First observedfidelis_query
    • First observedfidelis_recall

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation3/5

fidelis_recall and fidelis_query both retrieve memories and are distinguished mainly by filtering/LLM behavior, so an agent could easily choose the wrong one. fidelis_orient adds another retrieval-like path, though fidelis_health is clearly distinct.

Naming Consistency4/5

All tools share the fidelis_ prefix and use consistent snake_case, making the style predictable. The minor deviation is that fidelis_health is a noun rather than an action verb like recall, query, or orient.

Tool Count5/5

Four tools is a tight, manageable surface for a focused memory-retrieval server. The count is well within the ideal range and each tool has a recognizable role.

Completeness2/5

The domain is memory, yet the toolset only reads, queries, checks health, and orients. There are no operations to store, update, or delete memories, so agents cannot persist new information or correct stale memories.

Maintenance

ActivityActive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    Provides persistent local memory functionality for AI assistants, enabling them to store, retrieve, and search contextual information across conversations with SQLite-based full-text search. All data stays private on your machine while dramatically improving context retention and personalized assistance.
    3
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides local-first, cross-session memory for Claude Code, enabling semantic search across past sessions to retrieve procedures, decisions, or answers without exposing secrets.
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides fully local long-term memory for AI agents by enabling semantic search over notes and session logs using Ollama embeddings, with no external APIs or databases.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides coding agents with persistent local memory by exposing tools to save, search, and retrieve decisions, bugs, and context as Markdown with hybrid keyword and semantic search.
    4
    MIT