glean-company-docs
Halcyon Docs Chatbot
Beantwortung von Fragen auf Basis eines lokalen Dokumentkorpus, aufgebaut auf Gleans Indexing-, Search- und Chat-APIs und bereitgestellt als einzelnes MCP-Tool.
Es gibt keine Weboberfläche. Die Chat-Oberfläche ist ein MCP-Client – Cursor, Claude Desktop oder ein anderer.
So funktioniert es
data/Halcyon Shared Drive/ -> Indexing API -> Search API -> Chat API -> {answer, sources, diagnostics}Search läuft vor Chat, und Chat ruft nie selbst ab: Passagen werden explizit abgerufen und übergeben. Wenn nichts die Relevanzschwelle überschreitet (Begriffsüberlappung mit der Frage), lautet die Antwort ehrlich „keine indizierten Inhalte gefunden“ und Chat wird nie aufgerufen.
Related MCP server: faq-rag
Einrichtung
Erfordert Poetry und Python 3.12+.
poetry config virtualenvs.in-project true # keeps the venv at ./.venv
poetry install
cp .env.example .env # then fill in the tokens.env ist in gitignored. Committe niemals echte Tokens.
Variable | Verwendet von | Hinweise |
| beide | SDK baut |
| nur Indexing | wird auf dem Query-Pfad nie geladen |
| Search + Chat | Scope Chat/Search, Typ Global |
| beide | gemeinsame Sandbox, daher namespaced dies Doc-IDs |
| nur Indexing | Korpus-Root ( |
| Search + Chat | E-Mail, als die agiert wird; für Global-Tokens erforderlich |
Optional, mit Standardwerten: GLEAN_DOC_ID_PREFIX (halcyon), GLEAN_TOP_K (5),
GLEAN_MAX_SNIPPET_SIZE (2000), GLEAN_MIN_TERM_OVERLAP (0.30),
GLEAN_CHAT_TIMEOUT_MS (60000).
Die beiden Tokens sind strukturell getrennt: Settings.for_indexing() liest
GLEAN_INDEXING_TOKEN, Settings.for_query() liest GLEAN_CLIENT_TOKEN und
fasst die Indexierungsvariable nie an.
Verwendung
poetry run python -m glean_chat_bot # the MCP server, on stdio
poetry run glean-index --dry-run # extract and report, send nothing
poetry run glean-index # extract and bulk-push
poetry run glean-index --process-now # ask Glean to process immediately (1 per 3h)
poetry run pytest # contract tests: no network, no tokens
poetry run pytest -m live # the eval set, against real GleanFüge -v für Debug-Logging hinzu.
Indexierung ist asynchron: glean-index kehrt zurück, sobald Glean die Dokumente
akzeptiert hat, Minuten bevor sie durchsuchbar werden. Bestätige die vollständige Abdeckung in der
Glean-Admin-Konsole, bevor du dich auf die Antworten verlässt.
MCP-Client-Konfiguration
{
"mcpServers": {
"glean-company-docs": {
"command": "/absolute/path/to/glean-chat-bot/.venv/bin/python",
"args": ["-m", "glean_chat_bot"],
"env": {
"GLEAN_INSTANCE": "support-lab",
"GLEAN_CLIENT_TOKEN": "...",
"GLEAN_ACT_AS": "you@example.com",
"GLEAN_DATASOURCE": "interviewds3"
}
}
}
}Ein Tool, ask_company_docs(question, top_k=None, include_citations=True) -> dict,
das {answer, sources, diagnostics} zurückgibt. diagnostics berichtet, was durchsucht
wurde und was zurückkam, damit das aufrufende Modell „nichts gefunden“ von „meine Formulierung
hat gefehlt“ unterscheiden und entsprechend erneut versuchen kann.
Aufbau
glean_chat_bot/
__main__.py `python -m glean_chat_bot`, the MCP server: one tool over query.ask.ask()
client.py indexing and query client factories, the ActAs header
extraction.py one adapter per file type, path signals, walk
indexing.py the whole write path, behind the `glean-index` command
models.py Passage, Source, Answer, ExtractedDoc (pydantic)
query/ search.py (search -> Passage, the relevance floor)
chat.py (chat -> answer + resolved citations)
ask.py (ask() — the single orchestration function)
utils/ config.py (env loading, one Settings, two constructors)
logging.py (log format, timing wrapper on every Glean call)
data/ the corpus
docs/ extraction notes
tests/ test_contract.py (the invariants, offline)
eval_cases.py + test_eval_live.py (the eval set, `-m live`)Poetry für Abhängigkeiten und Paketierung, Ruff für Linting und Formatierung. Führe
poetry run ruff check . und poetry run ruff format . vor dem Committen aus.
Noch nicht gebaut
Gruppen- und Benutzerberechtigungen, Abteilungsfilter, Frische-Anmerkungen, ein Inhalts-Hash-Manifest, Query-Umschreibung und adaptives Wiederholen, Streaming, Konversationsspeicher, Wiederholung und Backoff, Docker, CI.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables grounding AI responses in a local document corpus by exposing MCP tools to list, search, and summarize documents, and generating answers using OpenAI.
- FlicenseNot gradedqualityDmaintenanceEnables answering natural-language questions from FAQ documents using vector search and LLM generation via an MCP tool.
- AlicenseNot gradedqualityBmaintenanceEnables hybrid document search (BM25 and dense) over a configurable corpus via MCP tools, returning passages and sources for AI agents to cite in answers.MIT
- AlicenseNot gradedqualityBmaintenanceEnables document-based Q&A with multi-modal RAG, hybrid retrieval, knowledge graph reasoning, and multi-agent orchestration via MCP tools.4MIT
Related MCP Connectors
Query any docs site via MCP. Submit a URL, ask questions, get cited answers.
Google AI Overview answers and cited sources via the Apify Google AI Overview API, hosted MCP.
Your company's brain for AI agents. Cited, permission-aware knowledge across every system.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/cmacdonald0514/glean-chat-bot'
If you have feedback or need assistance with the MCP directory API, please join our Discord server