Skip to main content
Glama
leoge007

agent-fact-system

by leoge007
README.md
# Agent Fact System

[中文](README.zh-CN.md) | English

<p align="center">
  <img src="assets/afs-hero.jpg" alt="Agent Fact System" width="768">
</p>

**Lightweight. Fast. Unconstrained.**

Agent Fact System, or AFS, is a local-first knowledge system for reasoning agents. It keeps facts, evidence, revisions, documents, retrieval indexes, and audit history outside the model context, then exposes them through a compact CLI and a thin stdio MCP server.

AFS is purpose-built for the working style of frontier models such as **DeepSeek V4 Pro**. These models can plan, call tools, compare evidence, and revise conclusions. They work better with a small deterministic knowledge surface than with a large application framework surrounding them.

The design target is model-class specific and provider neutral. AFS does not contain a DeepSeek-only adapter and does not require a particular chat runtime. DeepSeek V4 Pro, another frontier model, or a local Agent can use the same contracts.

## Why AFS exists

A strong model can reason across a difficult task. It still needs a reliable answer to a simpler question. What is true now, where did that claim come from, and what changed since the last run?

AFS gives the Agent one factual authority with explicit evidence and readback. The model remains free to reason. The knowledge layer stays small, inspectable, and recoverable.

## What is implemented

- A transactional SQLite canonical store for facts and practices
- Evidence-bound proposals, revisions, status transitions, and idempotent writes
- Append-only Timeline events and revision history
- A portable Markdown document store with stable slugs and source line locators
- Structured, full-text, temporal, vector, and hybrid retrieval
- Tags, links, backlinks, and partial document resolution
- Extractive answers with citations
- A fixed allowlist of 14 Agent-facing MCP tools
- A thin JSON CLI for local operation and administration
- Rebuildable local and vector indexes
- Obsidian-compatible Agent projections and a narrow command inbox
- Verified backup, restore, health, and drift checks
- Fail-closed sensitivity rules that keep restricted content out of FTS, projections, exports, and remote embeddings
- Optional one-time GBrain migration and private cold-archive support

## Why it fits DeepSeek V4 Pro class models

AFS gives a high-end Agent a compact set of operations instead of another orchestration framework.

- **Small tool surface**  The MCP interface is capped at 14 tools.
- **Deterministic contracts**  JSON schemas, idempotency keys, revision checks, and canonical readback make tool results verifiable.
- **Evidence stays close**  Claims carry source locators, excerpts, hashes, and Timeline history.
- **Context stays lean**  The Agent retrieves the exact record or document span it needs.
- **Reasoning stays free**  AFS governs stored knowledge without prescribing how the model plans or thinks.
- **The runtime stays light**  Python, SQLite, three direct runtime dependencies, and no mandatory daemon.

This is what “unconstrained” means here. AFS does not try to become the Agent, the planner, or the application shell. It provides durable knowledge and gets out of the way.

## Architecture

```text
Agent or MCP client
        |
        v
  CLI / stdio MCP
        |
        +-------------------+
        |                   |
        v                   v
Canonical SQLite      Markdown documents
        |                   |
        +---------+---------+
                  |
                  v
       Rebuildable derived views
       FTS / Timeline / vectors / Vault
```

Canonical data remains authoritative. Full-text indexes, vector generations, and Vault projections can be rebuilt.

## Quick start

AFS requires Python 3.12 and [uv](https://docs.astral.sh/uv/).

```sh
git clone https://github.com/leoge007/agent-fact-system.git
cd agent-fact-system
uv sync --locked

export AFS_HOME="$HOME/.local/share/agent-fact-system"
install -d -m 0700 "$AFS_HOME"

uv run afs init --json
uv run afs doctor --json
```

Create a candidate fact through the public CLI.

```sh
uv run afs record propose --input - <<'JSON'
{"kind":"fact","subject":"projects/quickstart","claim":"AFS stores evidence-backed facts.","evidence":[{"locator":"inline:quickstart","excerpt":"AFS stores evidence-backed facts."}],"idempotency_key":"quickstart:propose:1"}
JSON
```

Build the local index and query it.

```sh
uv run afs index sync --json
uv run afs query 'evidence-backed facts' --mode fulltext --json
```

See [docs/quickstart.md](docs/quickstart.md) for MCP setup, documents, vectors, backup, and restore.

## MCP setup

Any stdio MCP client can launch AFS with an explicit data directory.

```json
{
  "command": "uv",
  "args": [
    "--directory",
    "/absolute/path/to/agent-fact-system",
    "run",
    "python",
    "-m",
    "afs.mcp"
  ],
  "env": {
    "AFS_HOME": "/absolute/path/to/afs-data"
  }
}
```

The repository also includes an Agent routing Skill at `skills/agent-fact-system/SKILL.md`.

## Retrieval and embeddings

Local structured, full-text, temporal, and document retrieval work without a network service. Vector and hybrid retrieval are optional.

The current remote embedding adapter uses SiliconFlow with `Qwen/Qwen3-Embedding-8B`. Only records marked `normal` are eligible for remote embedding. Restricted content fails closed before an HTTP request is made.

```sh
export SILICONFLOW_API_KEY='...'
uv run afs embedding preflight --json
uv run afs index vector --json
uv run afs query 'your question' --mode hybrid --json
```

## Dream: opt-in session extraction

Dream adds an explicit `afs dream run` workflow for a local Hermes session database.
It writes episode ledgers and review candidates to staging, not confirmed facts.

- Successful inputs are tracked by session and input fingerprint across runs;
  failed or budget-skipped inputs remain gaps.
- Per-session file locks and unique run directories avoid duplicate concurrent work.
- Weekly source helpers include successful sessions from partial runs and preserve
  the original producing run when a later run inherits its results.
- Historical completion backfill requires recorded input evidence; repair and
  backfill commands default to dry-run.
- The chat adapter supports explicit reasoning effort and bounded transport retries;
  SiliconFlow embedding read timeouts also receive bounded retries.

Configure `~/.config/agent-fact-system/dream.json` and the selected provider's
API-key environment variable before using `afs dream run --date YYYY-MM-DD`.
`afs dream run --date YYYY-MM-DD --dry-run` plans without a model call, but may
create staging directories. Real extraction sends conversation text to the chosen
provider and may incur costs. Review the source scope and budget first; bundled
cost rates are estimates, not a spending guarantee. Dream currently uses POSIX
file locks (macOS/Linux). No scheduler or personal runtime configuration is shipped.
The weekly selection API is `afs.dream.weekly_source`; this release does not ship
a hosted review service or automatically promote candidates.

## Deliberate boundaries

AFS currently targets Python 3.12. It does not run a daemon, scrape conversations automatically, promote model output into confirmed facts, or hide provider failures behind silent fallback. Owner-level mutations remain outside the normal MCP surface.

The remote embedding adapter is currently specific to SiliconFlow. The Agent model itself remains independent of that adapter.

## Verification

```sh
uv sync --locked
uv run pytest --ignore=tests/live
```

Live embedding tests require an explicit API key and network authorization.

## License

MIT

TDQS

B3.2/5.0

Scored across 14 tools

Disambiguation5/5

Each tool targets a distinct resource/action: fact_get, fact_search, and fact_propose cover fact reading and proposing; timeline, health, and stats are separate system views; the eight document_* tools each perform a unique operation on documents. There is only a very minor overlap between health and stats, but their descriptions clearly differentiate them.

Naming Consistency3/5

The fact_* and document_* tools use a prefix-plus-verb pattern (fact_get, document_put), but timeline, health, stats are bare nouns, and document_tags, document_links, and document_backlinks are noun-noun constructions. The mixed conventions are readable but not consistent enough for a high score.

Tool Count5/5

Fourteen tools is within the well-scoped range for a system that manages facts and optional document capabilities. Each tool has a clear purpose, and the count is reasonable for the domain.

Completeness4/5

The fact lifecycle covers get, search, and propose, which is appropriate for a canonical/adjudicated system. Document tools cover get, list, search, put, resolve, tags, links, and backlinks, but there is no explicit delete operation for documents or facts, which is a minor gap.

Maintenance

ActivityMaintained
ResponsivenessNo issues