citetrail
Citetrail
Локальная память с подтверждённым происхождением о том, что видел ваш браузер — каждое воспоминание несёт URL, заголовок и временную метку, откуда оно взялось.
Citetrail захватывает страницы, которые вы действительно читали, хранит их на вашем компьютере и делает их доступными для поиска — вами и вашими ИИ-агентами через MCP. Когда агент использует что-то, что он там нашёл, он может точно указать, откуда это взялось.
Статус: предрелизная версия. Перед установкой см. Статус проекта.
Лицензия: Apache-2.0
Локально по умолчанию. Никакого аккаунта, сервера и загрузок. Заблокированные страницы не сохраняются (fail closed).
Проблема, которую решает Citetrail
Вы открыли шесть вкладок, закрыли их, и теперь вашему агенту-программисту нужна информация из четвёртой вкладки. Ваши варианты сегодня: вставить её снова, позволить агенту заново искать в открытом интернете и надеяться, что он попадёт на ту же страницу, или принять ответ без источника.
История браузера знает, что вы посетили URL. Она не знает, что было на странице, и не может сообщить это вашему агенту. Citetrail устраняет этот разрыв:
История браузера | Citetrail |
Список URL | Содержимое, которое вы действительно читали, сохранённое |
Поиск по заголовку, приблизительно | Поиск по тому, что было сказано на странице |
Невидимо для ваших инструментов | Доступно для запросов агентам через MCP |
Нет понятия «почему это здесь» | Каждая запись несёт своё происхождение |
Всё подряд | Только разрешённые страницы; чёрный список блокирует по умолчанию |
Related MCP server: qsearch
Что здесь означает «подтверждённое происхождение»
Каждый сохранённый фрагмент хранит ограниченную ссылку: исходный URL, заголовок страницы, временную метку захвата и позицию на странице. При извлечении возвращается фрагмент и эта ссылка вместе — их нельзя разделить. Агент, отвечающий из Citetrail, всегда может сказать, откуда он это взял, а вы всегда можете открыть оригинал.
Если источник исчез, Citetrail сообщает, что источник исчез. Он не выдаёт фрагмент за актуальный, как будто он всё ещё доступен.
Быстрый старт
git clone https://github.com/anonb3ll/citetrail
cd citetrail
python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/citetrail init
# 2. Search the local store
.venv/bin/citetrail search "retry backoff"
# Optional: block a sensitive hostname before it can be stored
.venv/bin/citetrail block bank.example.test
# 3. Point an agent at the same local store over MCP
.venv/bin/citetrail mcp --stdioХранилище по умолчанию — ~/.local/share/citetrail. Установите CITETRAIL_STORE или передайте --store PATH, чтобы использовать другую локальную директорию. См. docs/extension.md, чтобы загрузить распакованный адаптер Chromium.
Документация
Руководство | Описание |
Индекс документации | |
Команды CLI и структура хранилища | |
Схема инструментов MCP и регистрация | |
Настройка расширения Chromium | |
Чёрный список и поведение fail-closed | |
Опциональная интеграция с Runroom |
Часто задаваемые вопросы
Как разрешить моему ИИ-агенту искать в истории браузера?
Запустите локальный MCP-сервер и зарегистрируйте его у своего агента. Агент обращается к Citetrail как к любому другому MCP-инструменту и получает фрагменты с прикреплёнными источниками. Он никогда не получает прямого доступа к вашему браузеру или профилю.
Где хранятся мои данные и загружается ли что-то?
На вашем компьютере, в локальной базе данных, которую вы можете удалить в любой момент. У Citetrail нет сервера, и он ничего не загружает. См. docs/privacy.md.
Как запретить захват моего банка, почты или рабочей интрасети?
Чёрный список. Он проверяется перед захватом и блокирует по умолчанию — если правила не могут быть оценены для страницы, эта страница не захватывается. Добавьте хост с помощью citetrail block bank.example.test. Захват только по белому списку отложен.
Может ли агент цитировать источник, который он на самом деле не читал?
Не из Citetrail. Ссылка путешествует вместе с фрагментом; не существует API, который возвращает текст без его происхождения.
Что произойдёт, если я офлайн или страница исчезла?
Поиск работает офлайн по уже захваченным данным. Если исходный URL недоступен, результаты помечаются как таковые, а не молча выдаются за актуальные. Недоступные и заблокированные по соображениям конфиденциальности состояния честно сообщаются, а не скрываются.
Это приложение для заметок или «второй мозг»?
Нет. Citetrail захватывает и извлекает; он не организует ваше мышление, не строит граф знаний и не требует от вас ничего поддерживать. Это инфраструктура для инструментов, которым нужно знать, что вы читали.
Работает ли это в любом браузере?
Расширение в первую очередь ориентировано на браузеры на основе Chromium. Нативный мост между расширением и локальным сервисом имеет реальные ограничения — см. docs/limitations.md.
Чем Citetrail не является
Не облачный сервис и не сервис синхронизации. Один компьютер — одно хранилище.
Не система PKM и не заметки.
Не клинический инструмент, не инструмент для благополучия и не отслеживания внимания. Он не делает никаких заявлений о вашем когнитивном состоянии.
Не парсер. Он захватывает страницы, которые вы сами посетили, по вашим правилам.
Не мобильное приложение.
См. docs/limitations.md и docs/private-exclusions.md.
Связанный проект
Runroom координирует передачу задач между ИИ-агентами и людьми с помощью контрольных точек проверки и аудиторского следа. Оба проекта независимы, и ни один не требует другого; опциональная интеграция показывает ссылку Citetrail, передаваемую в управляемую задачу Runroom.
Участие
Прочтите CONTRIBUTING.md и CODE_OF_CONDUCT.md. Сообщайте об уязвимостях конфиденциально — см. SECURITY.md.
Статус проекта
Предрелизная версия, до 1.0. Интерфейсы будут меняться. Citetrail опубликован, чтобы выяснить, нужен ли он другим людям — если вы попробуете, расскажите нам, что вы пытались вспомнить и получилось ли у вас.
Лицензия
Apache License 2.0. Авторские права 2026, участники Citetrail.
Available Tools
1 toolcitetrail_searchC
Search local captures with inseparable provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| source_state | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. 'Search' implies a read-only operation, but the description does not clarify what 'inseparable provenance' means, how results are returned, whether source_state affects behavior, or what happens when captures are unavailable or privacy-blocked. This is a minimal signal rather than transparent behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler or repetition. The core action and resource are front-loaded. It is appropriately concise, though the cryptic 'inseparable provenance' could have been replaced with more useful information without harming length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple with only two parameters, but there is no output schema, no annotations, and no parameter-level documentation. The description leaves critical details undefined: what 'local captures' are, what 'inseparable provenance' means, how query matching works, and what the response shape is. This is not enough for an agent to reliably invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about either parameter. 'query' and 'source_state' are completely undocumented, and the meaning of the source_state enum values is left entirely to inference. The description fails to compensate for the schema's lack of parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a search operation over 'local captures,' which identifies the tool's verb and resource. The phrase 'with inseparable provenance' adds a distinguishing quality, though it is jargon-heavy and not fully explained. With no sibling tools to differentiate from, this is clear enough for basic selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when searching local captures, giving some usage context. However, it provides no explicit guidance on when to prefer this tool over alternatives, no prerequisites, and no exclusions. The usage signal is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
citetrail_search
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap. The purpose of citetrail_search is singular and unambiguous.
The single tool name 'citetrail_search' follows a clear object-action pattern, and with only one tool there is no inconsistency to evaluate.
A single search tool feels insufficient for a server named 'citetrail', which implies a broader capture management lifecycle. One tool is too thin for the apparent scope of the domain.
The server only exposes search; there are no create, retrieve, update, delete, or list operations for captures. This leaves agents unable to ingest or manage captures, creating significant gaps and dead ends.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Agentic search over your Dewey document collections from any MCP-compatible client.
Personal knowledge MCP: capture bookmarks, notes & todos by chat; archive pages; search memory.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI tools to query a user's private, locally stored memories (notes, documents) with source citations, using the MCP protocol.12 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web searches with full content retrieval and multi-engine provenance, including trust scoring and local corpus persistence, via MCP integration.1 npm2Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables LLM agents to perform web search, scraping, and summarization outside their context window, receiving compact cited briefs via MCP while full pages are cached and viewable in a local web UI.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web research over MCP: search with a real browser, fetch JS-rendered pages, download PDFs and convert them to Markdown, then search and page through large documents using bounded previews so the context window isn't flooded.BSD Zero Clause