dmine-mcp
Archives and mines Discord channels through a logged-in browser session, providing tools to capture/backfill channels, search archived messages, export channels to markdown or JSONL, list known servers and channels, check per-channel watermarks, and manage archived media attachments.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@dmine-mcpsearch my archive for dinner plans"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
dmine
Bot-free Discord channel archiver & data miner. It drives a normal logged-in browser (Playwright + system Chrome) and reads whatever the account can see, exactly like a human scrolling. No bot token, no API key. Output: one shared SQLite archive plus the downloaded media files.
Unofficial and unaffiliated. This project is not affiliated with, endorsed by, or supported by Discord Inc. "Discord" is a trademark of Discord Inc. and is used here only to describe what the tool interacts with.
Read this before using it. dmine automates a user account. That is not what Discord's Terms of Service and API rules permit for bots — the supported path for programmatic access is the official API with a bot token. Automated account use can get an account limited or terminated, and the web client's internal behaviour can change at any time. Archive only content you are authorised to access, and respect the privacy of the people whose messages you are collecting. You are responsible for how you use this.
What it is not
Not a bot: it registers nothing with Discord's developer platform.
Not an API client: it makes no REST or gateway calls; it reads the rendered DOM.
Not a way to reach other people's private servers: it sees exactly what the logged-in account can already see in the client.
Related MCP server: Discord Message Finder MCP
Install
Requires Python 3.10+ and Google Chrome or Chromium installed on the system.
dmine drives your browser and starts it as a long-running background
daemon, so a single login survives across runs (that is what the persistent
profile and the --remote-debugging-port daemon are for). Chrome/Chromium is
looked up in the usual install locations and on PATH.
If none is found it falls back to Playwright's bundled Chromium
(playwright install chromium), which works but does not give you the
long-lived daemon — the browser then starts and stops with each command.
Installed from source — there is no package on PyPI:
git clone https://github.com/hschaefer/dmine.git && cd dmine
python3 -m venv .venv && .venv/bin/pip install -e .That gives you two entry points: the dmine CLI and dmine-mcp,
the MCP server (see Use as an MCP server).
One-time setup
dmine login
# prints a QR code (also written to --qr-out <path> if given)
# scan it with the Discord mobile app (Settings -> Scan QR code)
# the session token is injected into the browser profile automaticallyAlternatively, log in manually:
dmine browser start
# a browser window opens -> log into Discord once (2FA fine)The login persists in ~/.config/dmine/profile (never re-asked).
QR-code login (remote auth)
dmine login implements Discord's remote auth flow (the mechanism the
desktop app uses): it performs an RSA-OAEP key exchange with Discord's gateway,
renders https://discord.com/ra/<fingerprint> as a QR code, and — once you
scan it with the mobile app and confirm — exchanges the resulting ticket for a
regular session token, which is injected into the browser profile. No bot
token, no password, no API key.
This is an internal, undocumented flow. At the time of writing the web client ships the QR-login component internally but exposes no reachable UI toggle for it, so this library speaks to the remote-auth gateway directly. Treat it as best-effort: it is not covered by Discord's public API contract and can break without notice.
The flow is also available as a library:
from dmine.auth import RemoteAuth
from dmine.browser import DiscordBrowser
auth = RemoteAuth()
qr_url, png = auth.start() # handshake + QR code (PNG bytes)
token, user = auth.wait_for_token() # blocks until scanned+confirmed
browser = DiscordBrowser().start()
browser.set_session(token, user) # inject into the browser profile
browser.stop()set_session writes the session via CDP directly into the browser profile's
storage backend (JSON-encoded exactly as Discord Web's storage wrapper
expects). It deliberately avoids Discord's obfuscated localStorage hiding —
no unstable internal property names are used, so injection survives client
updates as long as Discord keeps the JSON-encoding storage convention.
Usage
dmine capture <channel-id-or-url> --full # one-time backfill
dmine capture <channel-id-or-url> # incremental since watermark
dmine search "keyword" [--channel <id>] [--server <id-or-name>]
dmine export <channel-id> --format md|jsonl [--out file]
dmine status [--channel <id>] # watermark per channel
dmine channels # known channelssearch is a case-insensitive substring match over message content, embeds and
author names — not a tokenised/ranked full-text index, and the searched columns
are unindexed, so a hit is found by scanning the archive. Fine for a personal
archive; expect it to get slower as the DB grows.
Use as an MCP server
dmine is also an MCP server (stdio
transport), so any MCP-capable client — Claude Desktop, Cursor, OpenCode,
Hermes, … — can use the archive as structured tools instead of shelling out to
the CLI. The MCP SDK is a normal dependency; there is nothing extra to install.
Register it by absolute path: MCP clients do not inherit your shell's
PATH or virtualenv.
{
"mcpServers": {
"dmine": {
"command": "/absolute/path/to/dmine/.venv/bin/dmine-mcp"
}
}
}python -m dmine.mcp_server is equivalent if you prefer to point the
client at an interpreter. DMINE_* variables (see below) can be passed
through the client's env block.
The tools fall into two groups:
Tool | Needs the browser session | What it does |
| no | most recent archived messages of a channel |
| no | substring search over content, embeds and author |
| no | total messages, per-channel watermark and counts |
| no | known servers (id, name, channel and message counts) |
| no | known channels, optionally of one server |
| no | archived attachments and embed images for a channel |
| no | export a channel to jsonl/markdown, sandboxed to |
| yes | archive a channel (incremental or full backfill) |
| yes | re-download media for already archived messages |
The read tools work against the SQLite archive alone and are safe to call at
any time. capture and remedia drive a real browser, so they need the
one-time setup above first; they are serialized by a lock and can run for a
long time, so they return a running job summary instead of blocking past a
short client timeout.
Architecture
dmine/— Python package (CLI + capture + storage). Harness-agnostic; the CLI is the shared core.dmine/mcp_server.py— thin MCP adapter (python -m dmine.mcp_server, stdio) so any MCP-capable client can call the same core as structured tools.archive.sqlite— single source of truth (messages,channels,servers). Message IDs are snowflakes →MAX(id)per channel is the incremental watermark.~/.config/dmine/media/<channel>/— downloaded attachments (CDN links expire, so media is captured at capture time).
Safety & limits
DOM-only; no network interception, no tokens, no API calls.
Scrolls with human-ish speed/jitter; stop conditions are convergence-based.
Only archives channels the account can actually view.
Threat model: the Chrome daemon exposes an unauthenticated CDP endpoint on 127.0.0.1 — any local process could attach and drive the logged-in Discord session. Keep the daemon profile on a trusted machine. See SECURITY.md.
Concurrency: all browser-driving operations take a non-blocking lock (
~/.config/dmine/capture.lock, PID-verified, stale locks are reclaimed). DB reads (search/status/recent/…) are never locked.The archive holds other people's messages. Treat it as personal data: don't commit it, don't redistribute it, don't attach exports or channel screenshots to issues.
.gitignorealready excludes the archive, the browser profile, media and QR codes.
Environment variables
Var | Default |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Tests
The suite is offline — no browser, no network, no Discord account, no existing archive. It uses synthetic IDs only.
.venv/bin/python -m pytest -qLicense
AGPL-3.0-or-later — see LICENSE. Running a modified version as a network service means section 13 requires you to offer its source to your users.
Contributing note
The archive path can be overridden with DMINE_DB; the package is imported as dmine.
Available Tools
9 toolscaptureA
Archive a Discord channel into the SQLite archive.
Incremental by default: only messages newer than the stored watermark (MAX message id per channel) are fetched. Use full=True for a one-time backfill. Requires the persistent Chrome daemon (auto-starts; the user must have logged into Discord in its profile at least once).
Long jobs run in the background: the call waits up to wait seconds
(default 45) and returns the full result if the capture finished in
time. If it is still running, it returns {"status": "running", ...}
immediately and the capture continues — call capture() again for the
same channel to fetch the final result, or watch progress via
status()/recent(). Never fire captures for several channels in
parallel: they queue on the global capture lock and serialize
automatically (capture() reports waiting_for_lock while queued).
Args: channel: channel ID or full discord.com/channels/... URL. full: backfill the entire history instead of incrementing. since: explicit watermark message ID (overrides the stored one). limit: hard cap on messages (test runs). no_media: skip attachment + embed-image downloads (CDN links expire!). refresh: re-extract ALL visible messages, retrofitting embed images onto rows captured before image extraction existed (combine with full=True for the whole channel). wait: seconds to wait synchronously for completion before returning a "running" summary (keep below your client's request timeout (many clients cap it at ~60 s). Returns: Dict with new_messages, seen, stop_reason, total_in_db — or a {"status": "running", "job": {...}} summary for long backfills.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | ||
| wait | No | ||
| limit | No | ||
| since | No | ||
| channel | Yes | ||
| refresh | No | ||
| no_media | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses the watermark-based incremental semantics, the persistent Chrome daemon dependency and login prerequisite, the background-job/lock-serialization model (waiting_for_lock, queued captures), the synchronous wait timeout caveat, and the running-status return shape. These are exactly the operational traits an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by mode semantics, prerequisites, concurrency behavior, and a structured Args/Returns block. Despite its length, every sentence adds decision-relevant information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, long-running, stateful tool with no output schema and zero schema coverage, the description covers purpose, modes, prerequisites, concurrency, timeouts, parameter meanings, and return values. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all 7 parameters, and it does: channel accepts an ID or full URL, full backfills history, since overrides the stored watermark, limit is a test-run cap, no_media skips downloads with an expiry warning, refresh retrofits embed images onto old rows, and wait carries a client-timeout caveat. Each parameter gains meaning well beyond its bare type/default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Archive a Discord channel into the SQLite archive') and immediately establishes scope (incremental by default vs. full backfill). This is clearly distinguishable from siblings like export, status, recent, and search, which handle different facets of the same archive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use full=True (one-time backfill) vs. the default incremental mode, names status()/recent() as alternatives for progress-watching, tells the agent to re-call capture() to fetch a finished job, and warns against parallel captures across channels. Both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channelsA
List channels known to the archive (optionally of one server).
Args: server: optional server ID or name to restrict to. Returns: Dict: {"channels": [ {id, server_id, server_name, name, last_message_id}, ... ]}.
| Name | Required | Description | Default |
|---|---|---|---|
| server | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full behavioral burden. It does disclose the return shape (a Dict with a channels list containing id, server_id, server_name, name, last_message_id) and implies a read-only list operation. However, it omits auth requirements, rate limits, pagination, and whether results are cached or live, leaving clear gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with its purpose, then uses compact Args and Returns sections. Every line serves a purpose: the parameter explanation and return shape are useful given the absent output schema. There is no redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by documenting the return dictionary structure. For a simple one-parameter list tool with no annotations, the description covers purpose, parameter semantics, and return shape. It still lacks operational details like auth, rate limits, or pagination, keeping it from a top score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must define the only parameter. It does so adequately: 'server: optional server ID or name to restrict to' explains that the parameter is optional, accepts an ID or name, and filters the result. It could add format examples or default behavior, but it compensates well for the missing schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'List channels known to the archive', giving a specific verb (List) and resource (channels). It also scopes the result with the optional server filter, so an agent can understand the operation without reading the schema. However, it does not explicitly differentiate this tool from sibling tools like servers or search, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description only explains the optional server parameter; it offers no guidance on when to choose channels over siblings such as servers, recent, or search. There are no prerequisites, exclusions, or alternatives mentioned, so an agent receives minimal routing help. This matches the rubric's 'no guidance' level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exportA
Export a channel from the archive to a file (jsonl or markdown).
The output path is sandboxed to the export directory (DMINE_EXPORT_DIR, default ~/.config/dmine/exports/) — absolute paths and '..' traversal are rejected.
Args: channel: channel ID or full URL. format: "jsonl" or "md". since: only messages with id > this message ID. out: filename relative to the export dir (default .). Returns: Dict with path, message_count.
| Name | Required | Description | Default |
|---|---|---|---|
| out | No | ||
| since | No | ||
| format | No | jsonl | |
| channel | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the sandboxing constraint (paths confined to DMINE_EXPORT_DIR, absolute paths and '..' rejected) and the return shape (dict with path and message_count). It stops short of covering failure modes (missing channel, empty archive) or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and format, then structured Args/Returns sections. Every line adds information; the only minor cost is the docstring-style formatting overhead.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no annotations and no output schema, the description covers the operation, path safety rules, all parameters, and the return shape. The main gap is guidance on when this tool is the right choice versus its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the Args section is doing all the work, and it defines each of the four parameters meaningfully: channel accepts an ID or full URL, format enumerates 'jsonl'/'md', since is an exclusive id lower bound, and out is a filename relative to the export dir with a documented default. This meaningfully compensates for the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Export') and resource ('a channel from the archive') plus the output formats (jsonl or markdown). It is clearly distinguishable from siblings like status, capture, recent, search, and channels, which are non-file-producing operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to choose export over siblings such as recent, search, or capture, nor does it state any exclusions or preconditions for use. The purpose is inferable, but no routing guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mediaA
List attachments AND embed images archived for a channel.
Downloaded files live under DMINE_MEDIA (default ~/.config/dmine/media//, embed images in /embeds/). Entries without local_path only carry the CDN URL, which may already be expired. Use remedia() to (re-)download while the links are valid.
Args: channel: channel ID or full URL. limit: max messages to scan (default 200). message_since: only messages with id > this message ID. Returns: Dict: {"media_dir": str, "count": int, "attachments": [ ... ]} with per-entry kind ("attachment"|"embed_image"), message_id, ts, author, filename, downloaded, local_path, url.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| channel | Yes | ||
| message_since | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful traits: where files are stored under DMINE_MEDIA, that some entries are URL-only and may already be expired, and how to re-fetch them. It omits auth/permission requirements and any rate/scope caveats, keeping it just shy of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then storage/expiry context, then a clean Args/Returns block. The structured layout is efficient, though it is slightly verbose for a simple listing tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Documents all arguments and describes the return dict shape (media_dir, count, attachments with per-entry fields), which compensates for the absent output schema. Only missing element is any note on permissions or error behavior, but the definition is otherwise complete for a read-only listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: channel is a 'channel ID or full URL', limit is 'max messages to scan (default 200)', and message_since restricts to 'messages with id > this message ID'. All three params gain clear meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and two distinct resource types (attachments AND embed images) scoped to a channel. This distinguishes it from siblings like remedia, which it explicitly references for downloading, so an agent can route confidently.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly points to remedia() as the alternative for (re-)downloading while links are valid, and notes that entries without local_path only hold an expirable CDN URL. That gives strong contextual guidance, though it stops short of explicit when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recentA
Read the most recent archived messages of a channel (no browser needed).
Args: channel: channel ID or full URL. since: only messages with id > this message ID (e.g. the watermark from status() before a capture, to see what's new). limit: max messages (default 100). include_embeds: include embed text (tweet payloads live there). Returns: Dict: {"messages": [ {id, ts, author, content, embeds, attachments}, ... ]}.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| since | No | ||
| channel | Yes | ||
| include_embeds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It implies read-only ('Read', 'archived') and adds useful context that it needs no browser, but it does not disclose auth/permission requirements or rate limits, and errors are unaddressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, then clean Args/Returns sections. Nearly every line earns its place, though the Returns block partially restates what the field names already imply.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description documents the returned dict shape and all four parameters well. For a simple read tool this is largely complete, missing only auth/error behavior and pagination nuance beyond limit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it defines channel as ID or full URL, explains 'since' semantics with a concrete example, states limit is a max, and explains include_embeds carries tweet payloads. This adds real meaning beyond the bare schema titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Read the most recent archived messages of a channel'. The parenthetical '(no browser needed)' meaningfully scopes it against browser-dependent siblings like capture. It does not explicitly name a sibling it differs from (e.g. search), so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated. The 'since' arg notes it takes 'the watermark from status() before a capture, to see what's new', which hints at a workflow, but there is no explicit when-to-use-this-vs-search-or-capture guidance or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remediaA
Re-download media (attachments + embed images) for ALREADY archived messages and persist local paths (idempotent, skips existing files). Also retries entries whose earlier download failed (no local file).
Like capture(), long runs continue in the background: the call returns
a {"status": "running", ...} summary after wait seconds at the latest
and the download continues — call remedia() again for the same channel
to fetch the final result.
Args: channel: channel ID or full URL. limit: only the N most recent messages (None = all). wait: seconds to wait synchronously before returning a running summary. Returns: Dict: {"channel_id", "messages_with_media", "media_entries"}.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | ||
| limit | No | ||
| channel | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does a good job: idempotency ('skips existing files'), retry semantics for failed downloads, and the async continuation contract after `wait` seconds, including re-invocation to collect the final result. It omits auth/permission requirements and any rate-limit or failure-mode detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then the async contract, then Args/Returns. The Args section duplicates schema fields, but that duplication is warranted given the 0% schema description coverage. No wasted filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description supplies the return dict keys ('channel_id', 'messages_with_media', 'media_entries') and explains the running-status case. Complete for correct invocation, though permissions and progress/pagination behavior for very large channels are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must define all three parameters, and it does: channel accepts an ID or full URL, limit takes the N most recent messages (None = all), and wait is the synchronous seconds before a running summary is returned. This meaningfully exceeds the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Re-download media (attachments + embed images)') scoped to 'ALREADY archived messages', which implicitly distinguishes it from a first-time capture. The contrast with capture() is reinforced by the background-execution comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear usage context: run it over already-archived messages, and it doubles as a retry path for entries whose prior download failed. It does not explicitly frame when NOT to use it (e.g. 'use capture() for new messages'), so the alternative selection is left partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Substring search over the archive (message content, embeds, author).
Without scope this spans every archived server. Use server= or channel= to restrict; every hit carries its server_id/server_name.
Args: query: search term (substring match; % and _ match literally). server: optional server ID or name (case-insensitive, unique substring ok) to restrict to. channel: optional channel ID/URL to restrict to (applies within the server scope, if any). limit: max hits (default 50). Returns: Dict: {"hits": [ {id, server_id, server_name, channel_id, author_name, ts, content}, ... ]}.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| server | No | ||
| channel | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that % and _ match literally (no wildcard semantics), that server matching is case-insensitive with unique-substring tolerance, that hit count is capped by limit (default 50), and what each hit contains. It does not explicitly label the operation read-only or describe truncation/ordering behavior when hits exceed limit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core behavior and scoping rule in two sentences before the Args/Returns blocks, which are well structured. The Returns section partially duplicates what the parameter list already implies, but nothing is wasted enough to be a problem.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description is complete enough for a correct call: it documents all arguments, the scoping model, and the return dictionary shape. The remaining gap is behavioral detail around result truncation and ordering when limit is hit, which an agent would need for large archives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: query is a substring match with literal % and _, server accepts ID or case-insensitive name/unique substring, channel accepts an ID or URL and applies within the server scope, and limit documents its default of 50. Every parameter gains meaning beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('substring search') and resource ('the archive'), and enumerates what is searched: message content, embeds, author. It is clear this is the discovery tool, distinct from listing-style siblings like recent or media, though it never names an alternative outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains that without scope the search spans every archived server, and directs the agent to server= or channel= to restrict, including how channel interacts with server scope. It stops short of stating when NOT to use it (e.g., versus recent) or any rate/result-ceiling guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serversB
List Discord servers known to the archive (id, name, channels, messages).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'List' implies a read-only operation and the parenthetical enumerates returned fields, but nothing is said about pagination, result size, ordering, auth requirements, or how 'known to the archive' is determined.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and the returned data folded into a parenthetical. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema listing tool, naming the returned fields (id, name, channels, messages) covers most of what an agent needs. The remaining gap is the lack of any indication of result set size or pagination behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema declares zero parameters, so there are no parameter semantics to document; per the rubric this yields a baseline of 4. The description correctly implies the tool is callable with no arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List Discord servers') and scopes it to 'known to the archive', which is clearer than a tautology. It does not explicitly distinguish itself from siblings such as 'channels' or 'search', both of which could plausibly return server-adjacent data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this tool versus alternatives like 'channels', 'search', or 'export'. An agent must infer from the name and sibling names alone that this is the top-level enumeration entry point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusB
Archive status: total messages, per-channel watermark and counts.
Args: channel: optional channel ID/URL to show only that channel. Returns: Dict with totals, per-channel rows (id, name, msgs, watermark) and active_captures (background capture/backfill jobs currently running or queued on the capture lock, if any).
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does usefully disclose the return structure and that active_captures reflects background jobs holding the capture lock, which is real behavioral context. However, it never states that this is a non-mutating read, nor any auth or cost characteristics, so it only partially covers the no-annotation gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose followed by compact Args/Returns blocks; every line is informative. The pseudo-header formatting is slightly heavier than needed for such a small surface, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by enumerating the returned dict fields (totals, per-channel rows with id/name/msgs/watermark, active_captures). For a single-parameter read tool this is close to complete, lacking only an explicit read-only assurance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single parameter, so the description must compensate, and it does: 'channel: optional channel ID/URL to show only that channel' explains both optionality and acceptable forms. That is meaningful beyond the schema's bare anyOf string/null.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific resource and scope: 'Archive status: total messages, per-channel watermark and counts.' That is a concrete verb+resource and clearly distinct from siblings like export, capture, recent, or search. It stops short of naming an alternative, so it is clear but not sibling-differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus siblings such as channels, servers, or recent. The channel argument implies a narrowing use case, but the agent is left to infer that this is the read-only inspection tool for archive progress.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
capture - First observed
channels - First observed
export - First observed
media - First observed
recent - First observed
remedia - First observed
search - First observed
servers - First observed
status
TDQS
Scored across 9 tools
Each tool targets a distinct action on the Discord archive: capture ingests messages, recent/search read them, export writes them to a file, media/remedia manage attachments, and servers/channels/status provide metadata. No two tools have overlapping primary purposes, so an agent can select unambiguously.
All names are lowercase single tokens without underscores or camelCase, which is a consistent format. However, the set mixes verbs (export, capture, search, remedia) and nouns/adjectives (status, servers, channels, media, recent), making the pattern slightly less predictable than a pure verb_noun scheme.
Nine tools is well within the ideal 3–15 range for a focused archive server. Each tool corresponds to a clear capability (ingest, monitor, read, search, export, navigate, media management) and none feels redundant.
The set covers the core archive lifecycle: capture, status, recent, search, export, channel/server listing, and media listing/re-download. Minor gaps exist, such as no delete/prune operation for messages or channels and no way to cancel a running background capture, but these are workaround-able or outside typical agent workflows.
Maintenance
Related MCP Connectors
Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.
Official remote MCP server for Archivist AI TTRPG campaign memory: characters, sessions, and more.
Read-only MCP server for The Quiet Protocol's engines, benchmarks, proof, and business data.
Related MCP Servers
- AlicenseBqualityAmaintenanceMCP server that exposes any Telegram-Archive instance to LLMs, enabling message search, chat browsing, and access to archived Telegram history.757 npm5GPL 3.0
- FlicenseAqualityDmaintenanceRead-only MCP server for finding Discord messages. It enables searching guild messages, locating messages from jump URLs, and reading context around results.71-
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables searching Discord messages using Discord's native search API. Designed for use with Claude Code to find project decisions and communications in Discord.-
- AlicenseAqualityAmaintenanceLocal-first, read-only MCP server for searching and retrieving cited evidence from archived Telegram chats, including transcripts and media metadata.53MIT