citetrail
Citetrail
本地、有来源追溯的浏览器所见记忆——每次回忆都带有来源 URL、标题和时间戳。
Citetrail 捕获你实际阅读过的页面,将它们保存在你的机器上,并使其可搜索——供你自己以及你的 AI 代理通过 MCP 使用。当代理使用了在其中找到的内容时,它可以准确引用内容的来源。
状态: 预发布。安装前请参阅 项目状态。
许可证: Apache-2.0
默认本地化。 无需账户、无需服务器、无需上传。被阻止的页面默认拒绝。
Citetrail 解决的问题
你打开了六个标签页,关闭了它们,现在你的编码代理需要第四个标签页里的内容。你目前的选择是:重新粘贴一次,让代理重新搜索开放网络并希望它落在同一页面上,或者接受一个没有来源的答案。
浏览器历史知道你访问过某个 URL。它不知道页面说了什么,也无法告诉你的代理。Citetrail 弥补了这一差距:
浏览器历史 | Citetrail |
URL 列表 | 你实际阅读过的内容,已捕获 |
大致按标题搜索 | 按页面内容搜索 |
对你的工具不可见 | 可被 MCP 上的代理查询 |
没有“为什么在这里”的概念 | 每条记录都带有其来源 |
一切内容,不加区分 | 仅允许的页面;阻止列表默认拒绝 |
Related MCP server: qsearch
这里“有来源追溯”的含义
每个存储的片段都保留一个有界引用:来源 URL、页面标题、捕获时间戳以及页面内的位置。回忆会同时返回片段和该引用——它们无法被分离。从 Citetrail 获得答案的代理总能说明内容来自哪里,而你总能打开原始页面。
如果来源已不存在,Citetrail 会说明来源已不存在。它不会悄悄地把片段当作仍然有效的内容来提供。
快速开始
git clone https://github.com/anonb3ll/citetrail
cd citetrail
python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/citetrail init
# 2. Search the local store
.venv/bin/citetrail search "retry backoff"
# Optional: block a sensitive hostname before it can be stored
.venv/bin/citetrail block bank.example.test
# 3. Point an agent at the same local store over MCP
.venv/bin/citetrail mcp --stdio默认存储位置是 ~/.local/share/citetrail。设置 CITETRAIL_STORE 或传递 --store PATH 以使用不同的本地目录。参见 docs/extension.md 以加载未打包的 Chromium 适配器。
文档
指南 | 描述 |
文档索引 | |
CLI 命令与存储布局 | |
MCP 工具模式与注册 | |
Chromium 扩展设置 | |
阻止列表与默认拒绝行为 | |
可选的 Runroom 集成 |
常见问题
如何让我的 AI 代理搜索我的浏览历史?
运行本地 MCP 服务器并将其注册到你的代理。代理像使用任何其他 MCP 工具一样查询 Citetrail,并接收带有来源附件的片段。它永远不会获得对你浏览器或配置文件的原始访问权限。
我的数据存储在哪里,会有任何上传吗?
存储在你的机器上的本地数据库中,你可以随时删除。Citetrail 没有服务器,也不执行任何上传。参见 docs/privacy.md。
如何阻止它捕获我的银行、电子邮件或工作内网?
使用阻止列表。它在捕获之前进行检查,并且默认拒绝——如果无法评估某个页面的规则,该页面就不会被捕获。使用 citetrail block bank.example.test 添加主机。仅允许列表的捕获功能暂缓。
代理能否引用它并未实际阅读过的来源?
不能从 Citetrail 做到。引用与片段一起传递;没有任何 API 会返回不带来源的文本。
当我离线或页面不存在时会发生什么?
回忆可以离线针对你已捕获的内容工作。如果原始 URL 无法访问,结果会被标记为不可访问,而不是被静默地呈现为当前内容。不可用和隐私阻止状态会被如实报告,而不是隐藏。
这是一个笔记应用还是“第二大脑”?
不是。Citetrail 只负责捕获和回忆;它不组织你的思维、构建知识图谱,也不要求你维护任何东西。它是为需要知道你阅读内容的工具提供的基础设施。
它在任何浏览器中都能工作吗?
扩展首先面向基于 Chromium 的浏览器。扩展与本地服务之间的原生桥接存在实际限制——参见 docs/limitations.md。
Citetrail 不是什么
不是托管服务,也不是同步服务。一台机器,一个存储。
不是 PKM 或笔记系统。
不是临床、健康或注意力追踪工具。它对你的认知不做任何断言。
不是爬虫。它只在你自己的规则下捕获你自己访问过的页面。
不是移动应用。
参见 docs/limitations.md 和 docs/private-exclusions.md。
相关项目
Runroom 通过审查门禁和审计跟踪来协调 AI 代理与人类之间的交接。这两个项目相互独立,互不依赖;一个可选的集成展示了 Citetrail 引用如何为受治理的 Runroom 任务提供输入。
贡献
阅读 CONTRIBUTING.md 和 CODE_OF_CONDUCT.md。请私下报告漏洞——参见 SECURITY.md。
项目状态
预发布,pre-1.0。接口将会变化。Citetrail 发布是为了了解其他人是否需要它——如果你尝试了,请告诉我们你当时想回忆什么,以及你是否找到了。
许可证
Apache License 2.0。版权所有 2026 The Citetrail Contributors。
Available Tools
1 toolcitetrail_searchC
Search local captures with inseparable provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| source_state | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. 'Search' implies a read-only operation, but the description does not clarify what 'inseparable provenance' means, how results are returned, whether source_state affects behavior, or what happens when captures are unavailable or privacy-blocked. This is a minimal signal rather than transparent behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler or repetition. The core action and resource are front-loaded. It is appropriately concise, though the cryptic 'inseparable provenance' could have been replaced with more useful information without harming length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple with only two parameters, but there is no output schema, no annotations, and no parameter-level documentation. The description leaves critical details undefined: what 'local captures' are, what 'inseparable provenance' means, how query matching works, and what the response shape is. This is not enough for an agent to reliably invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about either parameter. 'query' and 'source_state' are completely undocumented, and the meaning of the source_state enum values is left entirely to inference. The description fails to compensate for the schema's lack of parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a search operation over 'local captures,' which identifies the tool's verb and resource. The phrase 'with inseparable provenance' adds a distinguishing quality, though it is jargon-heavy and not fully explained. With no sibling tools to differentiate from, this is clear enough for basic selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when searching local captures, giving some usage context. However, it provides no explicit guidance on when to prefer this tool over alternatives, no prerequisites, and no exclusions. The usage signal is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
citetrail_search
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap. The purpose of citetrail_search is singular and unambiguous.
The single tool name 'citetrail_search' follows a clear object-action pattern, and with only one tool there is no inconsistency to evaluate.
A single search tool feels insufficient for a server named 'citetrail', which implies a broader capture management lifecycle. One tool is too thin for the apparent scope of the domain.
The server only exposes search; there are no create, retrieve, update, delete, or list operations for captures. This leaves agents unable to ingest or manage captures, creating significant gaps and dead ends.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Agentic search over your Dewey document collections from any MCP-compatible client.
Personal knowledge MCP: capture bookmarks, notes & todos by chat; archive pages; search memory.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI tools to query a user's private, locally stored memories (notes, documents) with source citations, using the MCP protocol.12 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web searches with full content retrieval and multi-engine provenance, including trust scoring and local corpus persistence, via MCP integration.1 npm2Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables LLM agents to perform web search, scraping, and summarization outside their context window, receiving compact cited briefs via MCP while full pages are cached and viewable in a local web UI.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web research over MCP: search with a real browser, fetch JS-rendered pages, download PDFs and convert them to Markdown, then search and page through large documents using bounded previews so the context window isn't flooded.BSD Zero Clause