Slackquery
Reads a Slackpipe-compatible canonical Slack archive stored in DuckDB as a read-only source and publishes the resulting FTS-plus-vector search artifact as an immutable, validated DuckDB database that the MCP server serves.
Enables hybrid search over a canonical Slack archive, exposing MCP tools to search messages with lexical, semantic, or fused (RRF) ranking, fetch individual messages with surrounding channel context, expand results into full chronological threads, and discover archived workspaces, channels, and coverage scopes.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Slackquerysearch #product for messages about the Q3 roadmap from the last month"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Slackquery
Private-by-design hybrid search for Slack archives
Documentation · Sister project: Slackpipe · Slackpipe wiki
Slackquery turns a canonical Slack archive in DuckDB into an immutable search artifact and serves it through Model Context Protocol (MCP). It combines BM25 lexical retrieval with exact cosine similarity and reciprocal rank fusion (RRF), while keeping ingestion data read-only and deployment under your control.
Why Slackquery
Hybrid retrieval: exact terms through DuckDB FTS, semantic matches through local embeddings, query-aware routing, weighted RRF, exact-term boosts, and deterministic thread/channel diversity.
Context-rich documents: messages, bounded thread contexts, and safely extracted text, code, CSV, PDF, DOCX, and PPTX attachment chunks.
Immutable serving: build and validate a new artifact before atomically publishing it; readers never share the mutable pipeline database.
Local model support: switch between a batched PyTorch-compatible server and Ollama with one environment variable.
Resumable enrichment: content-addressed vectors, leases, retries, and durable checkpoints make embedding work incremental and idempotent.
Read-only MCP tools: bounded search, exact message lookup, thread expansion, scope discovery, health probes, and optional bearer authentication.
Dagster-native operations: assets, checks, and event-driven reconciliation sensors are included without coupling the MCP process to the writer.
Operational controls: Prometheus metrics, blocking Gold validation, retention, rollback, and state backup/restore commands.
Related MCP server: Slack MCP Server
Architecture
flowchart LR
C[(Canonical Slack DuckDB<br/>read-only)]
F[(Attachment tree<br/>read-only)]
P[Deterministic projection]
S[(Pipeline state<br/>documents + vectors)]
E{Embedding backend}
PT[PyTorch-compatible<br/>embedding server]
OL[Ollama]
B[Build + validate<br/>immutable artifact]
A[(Published DuckDB<br/>FTS + vectors)]
M[MCP Streamable HTTP]
U[Agents and clients]
C --> P --> S
F --> P
S --> E
E --> PT
E --> OL
PT --> S
OL --> S
S --> B --> A --> M --> UThe canonical database is attached READ_ONLY. Projection, embeddings, build
metadata, and publication state live in Slackquery-owned storage. The serving
process opens only the selected artifact.
Quickstart
Prerequisites
Python 3.11+
a Slackpipe-compatible DuckDB archive
an Ollama-compatible
/api/embedendpoint, using either Ollama itself or the supported PyTorch transportthe DuckDB FTS extension (installed automatically by the container image)
Local installation
git clone https://github.com/thomasmaerz/slackquery.git
cd slackquery
python -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[dev]'
cp .env.example .envEdit the ignored .env for your database, storage directories, embedding
endpoint, and model. Then run the pipeline:
slackquery project
slackquery embedding-status
slackquery embed
slackquery build
slackquery validate /srv/slackquery/artifacts/search-*.duckdb --checksum
slackquery publish /srv/slackquery/artifacts/search-*.duckdb
slackquery runConnect an MCP client to http://mcp-host:8181/mcp. Liveness and readiness are
available at /healthz and /readyz; Prometheus metrics are at /metrics.
Docker Compose
Copy .env.example to .env, set the host mount variables, then run:
docker compose up --build -d
curl -fsS http://localhost:8181/healthz
curl -fsS http://localhost:8181/readyzCompose starts the read-only MCP service. Run projection, embedding, build, and publication from a writer process or through the included Dagster definitions.
Embedding backends
Both transports use the same /api/embed request shape. Switching transport is
safe only when the model weights, native dimensions, prefixes, 512-dimensional
truncation, and L2 normalization are identical.
Slackquery includes an optional authenticated PyTorch server for CUDA hosts. See the PyTorch embedding server guide for GPU installation, same-host and remote layouts, API-key setup, firewall requirements, and OpenAI-compatible usage. It is sandbox software and must not be exposed to the public internet.
EMBEDDING_BACKEND=pytorch
PYTORCH_EMBEDDING_BASE_URL=http://embedding-host:11435
EMBEDDING_API_KEY=replace-with-the-server-key
OLLAMA_EMBEDDING_BASE_URL=http://ollama-host:11434
EMBEDDING_MODEL=nomic-embed-text:v1.5Switch to Ollama without changing source code:
EMBEDDING_BACKEND=ollamaModel digests are intentionally not hard-coded. For reproducible generations, read the digest from your server and pin it locally:
SLACKQUERY_EMBEDDING_MODEL_REVISION=sha256-digest-from-your-serverWithout a revision, Slackquery labels the generation unpinned. Pinning is
recommended for production because it detects silent model replacement.
MCP tools
Tool | Purpose |
| Search in |
| Fetch one document with bounded same-channel context. |
| Expand a result into a chronological thread. |
| Discover archived workspaces, channels, and coverage. |
Responses include stable source identity, component ranks, artifact identity, resolved filters, and opaque pagination cursors. Fused scores are ranking values, not probabilities.
Project status
Tier | Capability |
Stable | Projection, resumable embeddings, immutable builds, validation, atomic publication, lexical search, exact semantic search, hybrid RRF, MCP tools, health probes. |
Beta | Dagster integration, remote multi-client operation, PyTorch-compatible transport, operational benchmarks. |
Automated Gold | Message/thread/file projection, route-aware weighted RRF, exact boosts, diversity, structural validation, recovery controls, and Prometheus observability. |
Not claimed | Human-judged relevance quality. Exact vector scan remains the default until measured SLO evidence justifies an approximate index. |
Documentation
Development
pytest
ruff check .
mypySee CONTRIBUTING.md for development expectations. Slackquery is available under the MIT License.
This server cannot be deployed
Maintenance
Related MCP Connectors
Make your knowledge agent-ready. One MCP endpoint, 5 connectors, 3 search modes.
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
Reddit & X data for AI agents over MCP. Semantic search, hosted, no Reddit API.
Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables LLMs to access Slack's search functionality to retrieve users, channels, messages, and thread replies from a Slack workspace.4-
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables AI agents to search and retrieve messages within a Slack workspace using the Slack Web API. It supports specific channel filtering and includes built-in rate limit handling for efficient message discovery.52 npmMIT
- AlicenseBqualityDmaintenanceAn MCP server that enables searching for messages and listing channels within a Slack workspace. It provides tools to retrieve channel history and filter messages based on text matching and date ranges.352 npmMIT
- FlicenseAqualityDmaintenanceAn MCP server for semantic search and retrieval of indexed Slack messages stored in Qdrant using Cohere reranking via AWS Bedrock. It enables users to search through Slack history, retrieve full message threads, and access channel or user statistics through natural language.5-