agent-fact-system
# Agent Fact System
[中文](README.zh-CN.md) | English
<p align="center">
<img src="assets/afs-hero.jpg" alt="Agent Fact System" width="768">
</p>
**Lightweight. Fast. Unconstrained.**
Agent Fact System, or AFS, is a local-first knowledge system for reasoning agents. It keeps facts, evidence, revisions, documents, retrieval indexes, and audit history outside the model context, then exposes them through a compact CLI and a thin stdio MCP server.
AFS is purpose-built for the working style of frontier models such as **DeepSeek V4 Pro**. These models can plan, call tools, compare evidence, and revise conclusions. They work better with a small deterministic knowledge surface than with a large application framework surrounding them.
The design target is model-class specific and provider neutral. AFS does not contain a DeepSeek-only adapter and does not require a particular chat runtime. DeepSeek V4 Pro, another frontier model, or a local Agent can use the same contracts.
## Why AFS exists
A strong model can reason across a difficult task. It still needs a reliable answer to a simpler question. What is true now, where did that claim come from, and what changed since the last run?
AFS gives the Agent one factual authority with explicit evidence and readback. The model remains free to reason. The knowledge layer stays small, inspectable, and recoverable.
## What is implemented
- A transactional SQLite canonical store for facts and practices
- Evidence-bound proposals, revisions, status transitions, and idempotent writes
- Append-only Timeline events and revision history
- A portable Markdown document store with stable slugs and source line locators
- Structured, full-text, temporal, vector, and hybrid retrieval
- Tags, links, backlinks, and partial document resolution
- Extractive answers with citations
- A fixed allowlist of 14 Agent-facing MCP tools
- A thin JSON CLI for local operation and administration
- Rebuildable local and vector indexes
- Obsidian-compatible Agent projections and a narrow command inbox
- Verified backup, restore, health, and drift checks
- Fail-closed sensitivity rules that keep restricted content out of FTS, projections, exports, and remote embeddings
- Optional one-time GBrain migration and private cold-archive support
## Why it fits DeepSeek V4 Pro class models
AFS gives a high-end Agent a compact set of operations instead of another orchestration framework.
- **Small tool surface** The MCP interface is capped at 14 tools.
- **Deterministic contracts** JSON schemas, idempotency keys, revision checks, and canonical readback make tool results verifiable.
- **Evidence stays close** Claims carry source locators, excerpts, hashes, and Timeline history.
- **Context stays lean** The Agent retrieves the exact record or document span it needs.
- **Reasoning stays free** AFS governs stored knowledge without prescribing how the model plans or thinks.
- **The runtime stays light** Python, SQLite, three direct runtime dependencies, and no mandatory daemon.
This is what “unconstrained” means here. AFS does not try to become the Agent, the planner, or the application shell. It provides durable knowledge and gets out of the way.
## Architecture
```text
Agent or MCP client
|
v
CLI / stdio MCP
|
+-------------------+
| |
v v
Canonical SQLite Markdown documents
| |
+---------+---------+
|
v
Rebuildable derived views
FTS / Timeline / vectors / Vault
```
Canonical data remains authoritative. Full-text indexes, vector generations, and Vault projections can be rebuilt.
## Quick start
AFS requires Python 3.12 and [uv](https://docs.astral.sh/uv/).
```sh
git clone https://github.com/leoge007/agent-fact-system.git
cd agent-fact-system
uv sync --locked
export AFS_HOME="$HOME/.local/share/agent-fact-system"
install -d -m 0700 "$AFS_HOME"
uv run afs init --json
uv run afs doctor --json
```
Create a candidate fact through the public CLI.
```sh
uv run afs record propose --input - <<'JSON'
{"kind":"fact","subject":"projects/quickstart","claim":"AFS stores evidence-backed facts.","evidence":[{"locator":"inline:quickstart","excerpt":"AFS stores evidence-backed facts."}],"idempotency_key":"quickstart:propose:1"}
JSON
```
Build the local index and query it.
```sh
uv run afs index sync --json
uv run afs query 'evidence-backed facts' --mode fulltext --json
```
See [docs/quickstart.md](docs/quickstart.md) for MCP setup, documents, vectors, backup, and restore.
## MCP setup
Any stdio MCP client can launch AFS with an explicit data directory.
```json
{
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/agent-fact-system",
"run",
"python",
"-m",
"afs.mcp"
],
"env": {
"AFS_HOME": "/absolute/path/to/afs-data"
}
}
```
The repository also includes an Agent routing Skill at `skills/agent-fact-system/SKILL.md`.
## Retrieval and embeddings
Local structured, full-text, temporal, and document retrieval work without a network service. Vector and hybrid retrieval are optional.
The current remote embedding adapter uses SiliconFlow with `Qwen/Qwen3-Embedding-8B`. Only records marked `normal` are eligible for remote embedding. Restricted content fails closed before an HTTP request is made.
```sh
export SILICONFLOW_API_KEY='...'
uv run afs embedding preflight --json
uv run afs index vector --json
uv run afs query 'your question' --mode hybrid --json
```
## Dream: opt-in session extraction
Dream adds an explicit `afs dream run` workflow for a local Hermes session database.
It writes episode ledgers and review candidates to staging, not confirmed facts.
- Successful inputs are tracked by session and input fingerprint across runs;
failed or budget-skipped inputs remain gaps.
- Per-session file locks and unique run directories avoid duplicate concurrent work.
- Weekly source helpers include successful sessions from partial runs and preserve
the original producing run when a later run inherits its results.
- Historical completion backfill requires recorded input evidence; repair and
backfill commands default to dry-run.
- The chat adapter supports explicit reasoning effort and bounded transport retries;
SiliconFlow embedding read timeouts also receive bounded retries.
Configure `~/.config/agent-fact-system/dream.json` and the selected provider's
API-key environment variable before using `afs dream run --date YYYY-MM-DD`.
`afs dream run --date YYYY-MM-DD --dry-run` plans without a model call, but may
create staging directories. Real extraction sends conversation text to the chosen
provider and may incur costs. Review the source scope and budget first; bundled
cost rates are estimates, not a spending guarantee. Dream currently uses POSIX
file locks (macOS/Linux). No scheduler or personal runtime configuration is shipped.
The weekly selection API is `afs.dream.weekly_source`; this release does not ship
a hosted review service or automatically promote candidates.
## Deliberate boundaries
AFS currently targets Python 3.12. It does not run a daemon, scrape conversations automatically, promote model output into confirmed facts, or hide provider failures behind silent fallback. Owner-level mutations remain outside the normal MCP surface.
The remote embedding adapter is currently specific to SiliconFlow. The Agent model itself remains independent of that adapter.
## Verification
```sh
uv sync --locked
uv run pytest --ignore=tests/live
```
Live embedding tests require an explicit API key and network authorization.
## License
MIT
TDQS
Scored across 14 tools
Each tool targets a distinct resource/action: fact_get, fact_search, and fact_propose cover fact reading and proposing; timeline, health, and stats are separate system views; the eight document_* tools each perform a unique operation on documents. There is only a very minor overlap between health and stats, but their descriptions clearly differentiate them.
The fact_* and document_* tools use a prefix-plus-verb pattern (fact_get, document_put), but timeline, health, stats are bare nouns, and document_tags, document_links, and document_backlinks are noun-noun constructions. The mixed conventions are readable but not consistent enough for a high score.
Fourteen tools is within the well-scoped range for a system that manages facts and optional document capabilities. Each tool has a clear purpose, and the count is reasonable for the domain.
The fact lifecycle covers get, search, and propose, which is appropriate for a canonical/adjudicated system. Document tools cover get, list, search, put, resolve, tags, links, and backlinks, but there is no explicit delete operation for documents or facts, which is a minor gap.