citetrail
Citetrail
Lokaler, herkunftsgesicherter Speicher dessen, was Ihr Browser gesehen hat – jede Erinnerung trägt die URL, den Titel und den Zeitstempel, von dem sie stammt.
Citetrail erfasst die Seiten, die Sie tatsächlich gelesen haben, speichert sie auf Ihrem Rechner und macht sie durchsuchbar – für Sie und für Ihre KI-Agenten über MCP. Wenn ein Agent etwas verwendet, das er dort gefunden hat, kann er genau angeben, woher es stammt.
Status: Vorabversion. Siehe Projektstatus vor der Installation.
Lizenz: Apache-2.0
Standardmäßig lokal. Kein Konto, kein Server, kein Upload. Blockierte Seiten schlagen fehl.
Das Problem, das Citetrail löst
Sie lesen sechs Tabs, schließen sie, und jetzt braucht Ihr Coding-Agent das Ding aus Tab vier. Ihre Optionen heute: erneut einfügen, den Agenten das offene Web durchsuchen lassen und hoffen, dass er auf dieselbe Seite stößt, oder eine Antwort ohne Quelle akzeptieren.
Der Browserverlauf weiß, dass Sie eine URL besucht haben. Er weiß nicht, was die Seite gesagt hat, und er kann es Ihrem Agenten nicht mitteilen. Citetrail schließt diese Lücke:
Browserverlauf | Citetrail |
Eine Liste von URLs | Der Inhalt, den Sie tatsächlich gelesen haben, erfasst |
Suche nach Titel, grob | Suche nach dem, was die Seite sagte |
Unsichtbar für Ihre Werkzeuge | Abfragbar durch Agenten über MCP |
Kein Konzept von „Warum ist das hier?“ | Jeder Eintrag trägt seine Herkunft |
Alles, wahllos | Nur erlaubte Seiten; Blockliste schlägt fehl |
Related MCP server: qsearch
Was „herkunftsgesichert“ hier bedeutet
Jedes gespeicherte Fragment behält eine begrenzte Referenz: Quell-URL, Seitentitel, Erfassungszeitstempel und die Position innerhalb der Seite. Der Abruf gibt das Fragment und diese Referenz zusammen zurück – sie können nicht getrennt werden. Ein Agent, der aus Citetrail antwortet, kann immer sagen, woher er es hat, und Sie können immer das Original öffnen.
Wenn die Quelle nicht mehr existiert, sagt Citetrail, dass die Quelle nicht mehr existiert. Es liefert nicht stillschweigend ein Fragment, als wäre es noch live.
Schnellstart
git clone https://github.com/anonb3ll/citetrail
cd citetrail
python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/citetrail init
# 2. Search the local store
.venv/bin/citetrail search "retry backoff"
# Optional: block a sensitive hostname before it can be stored
.venv/bin/citetrail block bank.example.test
# 3. Point an agent at the same local store over MCP
.venv/bin/citetrail mcp --stdioDer Standardspeicher ist ~/.local/share/citetrail. Setzen Sie CITETRAIL_STORE oder übergeben Sie --store PATH, um ein anderes lokales Verzeichnis zu verwenden. Siehe docs/extension.md, um den entpackten Chromium-Adapter zu laden.
Dokumentation
Anleitung | Beschreibung |
Dokumentationsindex | |
CLI-Befehle und Speicherlayout | |
MCP-Werkzeugschema und Registrierung | |
Einrichtung der Chromium-Erweiterung | |
Blockliste und Fail-Closed-Verhalten | |
Optionale Runroom-Integration |
Häufig gestellte Fragen
Wie lasse ich meinen KI-Agenten meinen Browserverlauf durchsuchen?
Starten Sie den lokalen MCP-Server und registrieren Sie ihn bei Ihrem Agenten. Der Agent fragt Citetrail wie jedes andere MCP-Werkzeug ab und erhält Fragmente mit ihren Quellen. Er erhält niemals direkten Zugriff auf Ihren Browser oder Ihr Profil.
Wo werden meine Daten gespeichert, und wird etwas hochgeladen?
Auf Ihrem Rechner, in einer lokalen Datenbank, die Sie jederzeit löschen können. Citetrail hat keinen Server und führt keine Uploads durch. Siehe docs/privacy.md.
Wie verhindere ich, dass es meine Bank, meine E-Mails oder mein Firmenintranet erfasst?
Die Blockliste. Sie wird vor der Erfassung geprüft und schlägt fehl – wenn die Regeln für eine Seite nicht ausgewertet werden können, wird diese Seite nicht erfasst. Fügen Sie einen Host mit citetrail block bank.example.test hinzu. Die Erfassung nur mit Whitelist ist aufgeschoben.
Kann ein Agent eine Quelle zitieren, die er nicht tatsächlich gelesen hat?
Nicht von Citetrail. Die Referenz reist mit dem Fragment; es gibt keine API, die Text ohne seine Herkunft zurückgibt.
Was passiert, wenn ich offline bin oder eine Seite nicht mehr existiert?
Der Abruf funktioniert offline gegen das, was Sie bereits erfasst haben. Wenn die ursprüngliche URL nicht erreichbar ist, werden Ergebnisse als solche markiert, anstatt stillschweigend als aktuell präsentiert zu werden. Nicht verfügbare und datenschutzblockierte Zustände werden ehrlich gemeldet, nicht versteckt.
Ist das eine Notiz-App oder ein zweites Gehirn?
Nein. Citetrail erfasst und ruft ab; es organisiert nicht Ihr Denken, baut kein Wissensdiagramm auf und verlangt nicht, dass Sie etwas pflegen. Es ist die Infrastruktur für Werkzeuge, die wissen müssen, was Sie gelesen haben.
Funktioniert es in jedem Browser?
Die Erweiterung zielt zuerst auf Chromium-basierte Browser. Die native Brücke zwischen der Erweiterung und dem lokalen Dienst hat echte Grenzen – siehe docs/limitations.md.
Was Citetrail nicht ist
Kein gehosteter Dienst und kein Synchronisierungsdienst. Ein Rechner, ein Speicher.
Kein PKM- oder Notizsystem.
Kein klinisches, Wohlbefindens- oder Aufmerksamkeits-Tracking-Werkzeug. Es macht keine Aussage über Ihre Kognition.
Kein Scraper. Es erfasst Seiten, die Sie selbst besucht haben, unter Ihren Regeln.
Keine mobile App.
Siehe docs/limitations.md und docs/private-exclusions.md.
Verwandtes Projekt
Runroom koordiniert Übergaben zwischen KI-Agenten und Menschen mit Review-Gates und einem Prüfpfad. Die beiden Projekte sind unabhängig und keines erfordert das andere; eine optionale Integration zeigt eine Citetrail-Referenz, die eine verwaltete Runroom-Aufgabe speist.
Mitwirken
Lesen Sie CONTRIBUTING.md und CODE_OF_CONDUCT.md. Melden Sie Schwachstellen privat – siehe SECURITY.md.
Projektstatus
Vorabversion, vor 1.0. Schnittstellen werden sich ändern. Citetrail wird veröffentlicht, um herauszufinden, ob andere Menschen das brauchen – wenn Sie es ausprobieren, sagen Sie uns, was Sie abrufen wollten und ob Sie es bekommen haben.
Lizenz
Apache License 2.0. Copyright 2026 The Citetrail Contributors.
Available Tools
1 toolcitetrail_searchC
Search local captures with inseparable provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| source_state | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. 'Search' implies a read-only operation, but the description does not clarify what 'inseparable provenance' means, how results are returned, whether source_state affects behavior, or what happens when captures are unavailable or privacy-blocked. This is a minimal signal rather than transparent behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler or repetition. The core action and resource are front-loaded. It is appropriately concise, though the cryptic 'inseparable provenance' could have been replaced with more useful information without harming length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple with only two parameters, but there is no output schema, no annotations, and no parameter-level documentation. The description leaves critical details undefined: what 'local captures' are, what 'inseparable provenance' means, how query matching works, and what the response shape is. This is not enough for an agent to reliably invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about either parameter. 'query' and 'source_state' are completely undocumented, and the meaning of the source_state enum values is left entirely to inference. The description fails to compensate for the schema's lack of parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a search operation over 'local captures,' which identifies the tool's verb and resource. The phrase 'with inseparable provenance' adds a distinguishing quality, though it is jargon-heavy and not fully explained. With no sibling tools to differentiate from, this is clear enough for basic selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when searching local captures, giving some usage context. However, it provides no explicit guidance on when to prefer this tool over alternatives, no prerequisites, and no exclusions. The usage signal is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
citetrail_search
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap. The purpose of citetrail_search is singular and unambiguous.
The single tool name 'citetrail_search' follows a clear object-action pattern, and with only one tool there is no inconsistency to evaluate.
A single search tool feels insufficient for a server named 'citetrail', which implies a broader capture management lifecycle. One tool is too thin for the apparent scope of the domain.
The server only exposes search; there are no create, retrieve, update, delete, or list operations for captures. This leaves agents unable to ingest or manage captures, creating significant gaps and dead ends.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Agentic search over your Dewey document collections from any MCP-compatible client.
Personal knowledge MCP: capture bookmarks, notes & todos by chat; archive pages; search memory.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI tools to query a user's private, locally stored memories (notes, documents) with source citations, using the MCP protocol.12 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web searches with full content retrieval and multi-engine provenance, including trust scoring and local corpus persistence, via MCP integration.1 npm2Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables LLM agents to perform web search, scraping, and summarization outside their context window, receiving compact cited briefs via MCP while full pages are cached and viewable in a local web UI.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web research over MCP: search with a real browser, fetch JS-rendered pages, download PDFs and convert them to Markdown, then search and page through large documents using bounded previews so the context window isn't flooded.BSD Zero Clause