citetrail
Citetrail
브라우저가 본 내용을 저장하는 로컬 기반 출처 추적 메모리 — 모든 회상에는 그것이 나온 URL, 제목, 타임스탬프가 함께 담깁니다.
Citetrail은 실제로 읽은 페이지를 캡처하여 사용자의 기기 안에 보관하고, 사용자와 사용자의 AI 에이전트가 MCP를 통해 검색할 수 있게 합니다. 에이전트가 여기서 찾은 내용을 사용할 때는 정확히 어디서 왔는지 인용할 수 있습니다.
상태: 프리릴리즈. 설치 전에 프로젝트 상태를 확인하세요.
라이선스: Apache-2.0
기본적으로 로컬. 계정도, 서버도, 업로드도 없습니다. 차단된 페이지는 fail-closed 방식으로 처리됩니다.
Citetrail이 해결하는 문제
여섯 개의 탭을 읽고 모두 닫았는데, 코딩 에이전트가 네 번째 탭에 있던 내용을 필요로 합니다. 현재 사용 가능한 방법은: 다시 붙여넣기, 에이전트가 열린 웹을 다시 검색하여 같은 페이지에 도달하기를 바라기, 또는 출처 없이 답을 받아들이기입니다.
브라우저 기록은 URL을 방문했다는 사실만 알지, 그 페이지가 무슨 내용이었는지는 모르고 에이전트에게 알려줄 수도 없습니다. Citetrail은 이 격차를 메웁니다:
브라우저 기록 | Citetrail |
URL 목록 | 실제로 읽은 내용을 저장한 콘텐츠 |
제목으로 대략 검색 | 페이지가 말한 내용으로 검색 |
도구에서 접근 불가 | MCP를 통해 에이전트가 쿼리 가능 |
"왜 여기 있는지" 개념 없음 | 모든 항목에 출처 정보 포함 |
무차별적으로 모든 것 | 허용된 페이지만; 블록리스트는 실패 시 차단 |
Related MCP server: qsearch
여기서 "출처 추적"이 의미하는 것
저장된 모든 조각에는 제한된 참조 정보가 포함됩니다: 원본 URL, 페이지 제목, 캡처 타임스탬프, 페이지 내 위치. 회상은 조각과 그 참조 정보를 함께 반환하므로 분리할 수 없습니다. Citetrail로부터 답하는 에이전트는 항상 어디서 가져왔는지 말할 수 있고, 사용자는 항상 원본을 열 수 있습니다.
소스 페이지가 사라졌다면 Citetrail은 그 사실을 명시합니다. 마치 아직 살아있는 것처럼 조각을 조용히 제공하지 않습니다.
빠른 시작
git clone https://github.com/anonb3ll/citetrail
cd citetrail
python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/citetrail init
# 2. Search the local store
.venv/bin/citetrail search "retry backoff"
# Optional: block a sensitive hostname before it can be stored
.venv/bin/citetrail block bank.example.test
# 3. Point an agent at the same local store over MCP
.venv/bin/citetrail mcp --stdio기본 저장소는 ~/.local/share/citetrail입니다. 다른 로컬 디렉토리를 사용하려면 CITETRAIL_STORE를 설정하거나 --store PATH를 전달하세요. 확장 프로그램의 압축 해제된 Chromium 어뎁터를 로드하려면 docs/extension.md를 참조하세요.
문서
가이드 | 설명 |
문서 색인 | |
CLI 명령과 저장소 레이아웃 | |
MCP 도구 스키마와 등록 | |
Chromium 확장 프로그램 설치 | |
블록리스트와 실패 시 차단 동작 | |
선택적 Runroom 통합 |
자주 묻는 질문
내 AI 에이전트가 검색 기록을 검색하게 하려면 어떻게 하나요?
로컬 MCP 서버를 실행하고 에이전트에 등록하세요. 에이전트는 다른 MCP 도구와 마찬가지로 Citetrail에 쿼리하며 출처가 첨부된 조각을 받습니다. 에이전트는 브라우저나 프로필에 직접 접근할 수 없습니다.
내 데이터는 어디에 저장되며 업로드가 있나요?
내 기기에 있으며 언제든 삭제할 수 있는 로컬 데이터베이스에 저장됩니다. Citetrail에는 서버도 없고 업로드도 없습니다. docs/privacy.md 참조.
은행? 이메일? 회사 인트라넷 캡처를 멈추려면?
블록리스트를 사용하세요. 이것은 캡처 전에 검사되며 실패 시 차단됩니다. — 규칙을 평가할 수 없으면 해당 페이지는 캡처되지 않습니다. 호스트를 차단하려면 citetrail block bank.example.test을 실행하세요. 허용 목록 전용 캡처는 보류된 상태입니다.
에이전트가 실제로 읽지 않은 소스를 인용할 수 있나요?
Citetrail에서는 불가능합니다. 참조 정보는 조각과 함께 이동하며, 출처 없이 텍스트를 반환하는 API는 없습니다.
오프라인이거나 페이지가 사라지면 어떻게 되나요?
오프라인에서도 이미 캡처한 내용을 검색할 수 있습니다. 원본 URL에 접근할 수 없으면 결과는 그렇게 표시되며 현재의 것처럼 조용히 제시되지 않습니다. 접근 불가 상태와 개인정보 차단 상태는 숨기지 않고 정직하게 보고됩니다.
노트 앱이나 세컨드 브레인인가요?
아닙니다. Citetrail은 캡처와 회사에만 전념합니다. 사고를 조직화하거나 지식 그래프를 만들거나 유지 관리를 요구하지 않습니다. 읽은 내용을 알아야 하는 도구들을 위한 파이프라인입니다.
모든 브라우저에서 동작하나요?
확장 프로그램은 우선적으로 Chromium 기반 브라우저를 대상으로 합니다. 확장과 로컬 서비스 사이의 네이티브 브리지에 실제 제한이 있습니다 — docs/limitations.md 참조.
Citetrail이 아닌 것
호스팅 서비스도 동기화 서비스도 아닙니다. 하나의 기기, 하나의 저장소.
PKM이나 메모 시스템이 아닙니다.
임상, 웹니스, 또는 주의 추적 도구가 아닙니다. 관련된 어떤 인지적 주장도 하지 않습니다.
스크래마가 아닙니다. 직접 방문한 페이지를 귀하의 규칙 아래 캡처합니다.
모바일 앱이 아닙니다.
docs/limitations.md와 docs/private-exclusions.md 참조.
관련 프로젝트
Runroom은 검토 게이트와 감사 추적을 통해 AI 에이전트와 인간 간의 핸드오프를 조정합니다. 두 프로젝트는 독립적이며 서로가 필수인 것은 아닙니다. 선택적 통합을 통해 특성 Citetrail 참조가 거버닝된 Runroom 작업을 지원할 수 있습니다.
기여
CONTRIBUTING.md와 CODE_OF_CONDUCT.md를 읽어주세요. 취약점은 비공개로 보고하세요 — SECURITY.md 참조.
프로젝트 상태
프리릴리즈, 1.0 이전입니다. 인터페이스는 변경될 수 있습니다. Citetrail은 다른 사람들이 이런 것이 필요한지 확인하려는 목적으로 공개되었습니다. — 이걸 시도한다면 어떤 것을 기억하려다 어떤 결과를 얻었는지 알려주세요.
라이선스
Apache License 2.0 라이선스. Copyright 2026 The Citetrail Contributors.
Available Tools
1 toolcitetrail_searchC
Search local captures with inseparable provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| source_state | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. 'Search' implies a read-only operation, but the description does not clarify what 'inseparable provenance' means, how results are returned, whether source_state affects behavior, or what happens when captures are unavailable or privacy-blocked. This is a minimal signal rather than transparent behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler or repetition. The core action and resource are front-loaded. It is appropriately concise, though the cryptic 'inseparable provenance' could have been replaced with more useful information without harming length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple with only two parameters, but there is no output schema, no annotations, and no parameter-level documentation. The description leaves critical details undefined: what 'local captures' are, what 'inseparable provenance' means, how query matching works, and what the response shape is. This is not enough for an agent to reliably invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about either parameter. 'query' and 'source_state' are completely undocumented, and the meaning of the source_state enum values is left entirely to inference. The description fails to compensate for the schema's lack of parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a search operation over 'local captures,' which identifies the tool's verb and resource. The phrase 'with inseparable provenance' adds a distinguishing quality, though it is jargon-heavy and not fully explained. With no sibling tools to differentiate from, this is clear enough for basic selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when searching local captures, giving some usage context. However, it provides no explicit guidance on when to prefer this tool over alternatives, no prerequisites, and no exclusions. The usage signal is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
citetrail_search
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap. The purpose of citetrail_search is singular and unambiguous.
The single tool name 'citetrail_search' follows a clear object-action pattern, and with only one tool there is no inconsistency to evaluate.
A single search tool feels insufficient for a server named 'citetrail', which implies a broader capture management lifecycle. One tool is too thin for the apparent scope of the domain.
The server only exposes search; there are no create, retrieve, update, delete, or list operations for captures. This leaves agents unable to ingest or manage captures, creating significant gaps and dead ends.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Agentic search over your Dewey document collections from any MCP-compatible client.
Personal knowledge MCP: capture bookmarks, notes & todos by chat; archive pages; search memory.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI tools to query a user's private, locally stored memories (notes, documents) with source citations, using the MCP protocol.12 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web searches with full content retrieval and multi-engine provenance, including trust scoring and local corpus persistence, via MCP integration.1 npm2Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables LLM agents to perform web search, scraping, and summarization outside their context window, receiving compact cited briefs via MCP while full pages are cached and viewable in a local web UI.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web research over MCP: search with a real browser, fetch JS-rendered pages, download PDFs and convert them to Markdown, then search and page through large documents using bounded previews so the context window isn't flooded.BSD Zero Clause