media-hunter-mcp
Search and download media from e621, rule34, E-Hentai/ExHentai, and Pixiv (images, galleries, videos, ugoira) via MCP tools.
Search posts across sites with filters like query, rating, min_score, page, limit.
Get full metadata for a specific post by ID.
Download a single post by ID (with subdir, timeout options).
Search and batch download up to 50 files, reporting errors/skips.
Download directly from a work page URL with automatic site/ID detection.
Run self_check to verify credentials and API connectivity for all sites.
Supports E-Hentai table/inner site selection, original image download, proxies, retries, timeouts, and file integrity verification.
Provides tools for searching and downloading Pixiv works, including images, videos, and multi-image galleries, with support for filters like rating, limit, and minimum bookmarks, as well as author-based directory organization and download by post ID or page URL.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@media-hunter-mcpsearch e621 for fox images and download the top 5"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
media-hunter-mcp
搜索并下载 e621、rule34.xxx、E-Hentai / ExHentai、Pixiv 的作品。支持图片、视频、多图画廊及 Pixiv ugoira;提供 MCP 工具和独立命令行。
1.3.2 更新
MCP 工具统一使用
media_hunter_前缀,减少与其他服务重名;搜索与下载彻底分开,搜索只返回元数据。单作品和多作品下载统一为
media_hunter_download,使用post_ids数组并返回逐作品列表。去重后一个作品不限文件数,多个作品共用 50 文件预算。补全工具说明、参数约束和输出结构,明确表里站、原图、分组下载、额度、超时与重试规则,供模型根据任务自行决策。
从 1.3.1 升级
本版调整了 MCP 工具名称和按 ID 下载的参数、返回结构。 升级后重启 MCP 服务并刷新客户端工具列表;写死旧工具名的提示词或工作流需按下表迁移。已有 config.toml 可继续使用,E 站的表里站和原图配置保持有效。
1.3.1 工具 | 1.3.2 对应调用 |
|
|
|
|
|
|
| 先 |
|
|
|
|
旧 MCP 名称不再注册。曾使用开发版 media_hunter_download_post / media_hunter_download_posts 的调用也统一改为 media_hunter_download。命令行移除 download-search,改为先 search、再用 download-posts 下载选定 ID;download 和 download-url 保留。
Related MCP server: NovelAI MCP Server
1.3.1 更新
E 站支持每次调用选择表站/里站;默认使用里站,按链接下载时默认遵循链接域名。
下载工具新增
original=true原图选项,配置默认允许使用;页面图与原图分别保存和校验复用。表站和里站代理分别使用
mirror_base、exhentai_mirror_base;原图入口的登录或额度错误会停止后续下载。
从 1.3.0 升级后,保留自己的 config.toml,按下文检查 E 站开关与代理项,并重启 MCP 服务以刷新工具参数。已有配置中显式指定的表站选择仍然有效。
快速开始
以下使用 Python 3.11+ 和 uv。首次获取项目:
git clone https://github.com/baichenxw/media-hunter-mcp.git
cd media-hunter-mcp
uv sync --locked --extra animation已有项目时直接进入项目目录。首次配置时执行下面的命令,仅在文件不存在时复制,避免覆盖已有凭证:
if (-not (Test-Path -LiteralPath config.toml)) {
Copy-Item -LiteralPath config.example.toml -Destination config.toml
}编辑本地 config.toml,填写需要使用的站点凭证,并修改 [network].proxy:示例值 http://127.0.0.1:7897 只适用于本机该端口确有代理的情况;不使用代理时填写 proxy = ""。完成后检查:
uv run media-hunter checkcheck 会检查全部四个站点,未配置凭证的站点可能失败并导致退出码为 1;请查看 JSON 中各站的 ok 和错误信息。某站检查失败不妨碍调用其他已配置站点。
animation 提供 ugoira 转 GIF/MP4 所需的 FFmpeg。查找顺序为 [sites.pixiv].ffmpeg_path → 系统 PATH 中的 ffmpeg → imageio-ffmpeg 提供的程序。已有系统 FFmpeg 时可仅运行 uv sync --locked;只保存原始动图包则设置 ugoira_format = "zip"。下载命令中显式添加 --extra animation 可确保可选依赖已安装,详见 uv 可选依赖说明。
使用 pip 安装
也可以使用 Python 3.11+ 自带的 pip,无需 uv。在克隆或解压后的项目目录中运行,以下为 Windows PowerShell 示例:
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install ".[animation]"按上文复制并编辑 config.toml 后,使用同一虚拟环境检查和启动:
.\.venv\Scripts\media-hunter.exe --config "config.toml" check
.\.venv\Scripts\media-mcp.exeLinux/macOS 将上述 .\.venv\Scripts\ 替换为 ./.venv/bin/,并去掉程序名的 .exe。如果已有系统 FFmpeg,或只保存 ugoira ZIP,可以将安装目标 ".[animation]" 改为 .。
也可直接从 GitHub 的版本标签安装或升级(需要 Git;此命令替代上面的本地安装命令):
.\.venv\Scripts\python.exe -m pip install --upgrade "media-hunter-mcp[animation] @ git+https://github.com/baichenxw/media-hunter-mcp.git@v1.3.2"直接安装不会在当前目录生成配置模板,请从 v1.3.2 的 config.example.toml 保存模板后配置。这里使用 GitHub 源码安装,不依赖同名 PyPI 包。pip 根据 pyproject.toml 解析依赖,不读取 uv.lock;需要按锁文件安装时使用上面的 uv 方式。语法参见 pip 官方文档。
接入 MCP 客户端时,将下方配置中的 command 改为仅含虚拟环境内 media-mcp.exe 绝对路径的数组,例如 ["C:/Projects/media-hunter-mcp/.venv/Scripts/media-mcp.exe"],并保留指向实际配置文件的 MEDIA_HUNTER_CONFIG。独立命令行示例则用该环境中的 media-hunter 替代 uv run media-hunter;安装时选择过 [animation] 后,无需再传 --extra animation。
配置
读取顺序:MEDIA_HUNTER_CONFIG 环境变量 → 当前目录 config.toml → 源码项目目录 config.toml → ~/.config/media-hunter/config.toml。命令行的 --config 优先于以上规则。相对下载目录以配置文件所在目录为基准。
配置 | 用途 |
| 下载根目录,支持 |
| HTTP(S)/SOCKS5 代理;空字符串表示直连,不读取系统代理环境变量 |
| 单次 HTTP 网络操作的超时秒数 |
| 每个候选地址的最大尝试次数,1–10,包含首次请求 |
| 可选 |
| 必填 |
|
|
| 必填 |
每站可设置 request_interval、download_delay、download_concurrency(1–16)。每次网络尝试都限速;download_delay 控制待下载文件开始处理的间隔,download_concurrency 控制同一作品内部的并发数。同一服务进程内,同站点的下载调用会排队,不同站点可以同时下载;排队时间计入总超时。并发数和重试次数必须填写整数,布尔开关使用 TOML 的 true / false。
媒体文件的连接失败、HTTP 可重试错误及传输中断共用 [network].retries 次尝试,不会因两层重试而相乘。API 请求仍按每个候选地址分别计数。
mirror_base 是用户自行配置的可信反向代理。e621/rule34 的 API 在连接失败、429 或可重试 5xx 后尝试镜像;E-Hentai 的 mirror_base 仅用于表站,exhentai_mirror_base 仅用于里站,分别作为所选站点的主地址,不跨站自动回退,支持 https://example.com/eh 这样的路径前缀。Pixiv 可分别设置 API/OAuth 主地址和图片镜像。API 镜像不会自动套用到图片 CDN 或 OAuth 地址。认证请求可能经配置的反向代理发送,请只使用自己信任的地址。
如果旧配置用 mirror_base 指向里站代理,请把该项改名为 exhentai_mirror_base;表站代理继续使用 mirror_base。
Pixiv 令牌在内存中自动续期。刷新返回的新 refresh token 会在当前服务进程内使用;重新启动仍读取配置中的值。配置文件不自动改写。
MCP 接入
以 OpenCode 为例,下面是完整 JSON 示例;已有配置时将 media-hunter 条目合并到原来的 mcp 对象中。两处 C:/Projects/media-hunter-mcp 都需要替换为自己的项目绝对路径:
{
"mcp": {
"media-hunter": {
"type": "local",
"command": [
"uv", "run", "--project",
"C:/Projects/media-hunter-mcp",
"--locked", "--extra", "animation", "media-mcp"
],
"enabled": true,
"environment": {
"MEDIA_HUNTER_CONFIG": "C:/Projects/media-hunter-mcp/config.toml"
}
}
}
}其他 MCP 客户端使用相同的命令、参数和环境变量,配置结构依客户端而定。默认使用 stdio,标准输出只传输 MCP 协议。更新项目后重新启动 MCP 客户端或其服务进程。
也支持 MEDIA_HUNTER_TRANSPORT=http(Streamable HTTP,默认 /mcp 路径);默认监听 127.0.0.1:8787。可用 MEDIA_HUNTER_HOST、MEDIA_HUNTER_PORT 覆盖。sse 仅保留给旧客户端,新接入使用 stdio 或 Streamable HTTP。HTTP 模式未配置身份验证,应在受信任的本机环境使用。
1.3.2 使用 FastMCP 4 / MCP Python SDK 2,支持 MCP 2026-07-28 的 server/discover、按请求协商、resultType 及列表缓存字段,也兼容旧版 initialize 流程。传输和版本转换交给 SDK;业务代码不自行拼接协议消息。工具声明只读/写入行为、参数范围和输出 JSON Schema。工具执行失败使用 isError=true,同时保留 JSON 文本与 structuredContent,部分完成的文件清单不会丢失。列表缓存提示为 60 秒、private,不缓存下载调用结果。
下载过程提供排队、详情、解析、文件完成和合成阶段的进度。只有客户端请求进度通知时才发送;消息中的页数表示当前作品进度,协议数值是单调递增的工作事件计数,总量未知时不伪造百分比。是否展示进度取决于客户端。
工具
工具 | 使用场景与结果 |
| 只搜索;返回 |
| 读取已知 ID 的作品详情,返回 |
| 只下载明确提供的同站点 |
| 已有作品页面链接时直接下载,自动识别站点和 ID;不接受搜索页面或媒体直链 |
| 检查凭证与 API/首页连通性,每站最多 45 秒;逐站查看 |
例如先调用 media_hunter_search:
{"site": "pixiv", "query": "風景", "rating": "safe", "limit": 5}搜索到作品后可以直接展示结果;仅在需要保存媒体时调用下载工具。将返回结果的 data.posts[].id 原样作为字符串放入 media_hunter_download 的 post_ids 数组,单个作品也使用数组:
{"site": "pixiv", "post_ids": ["123"]}多个作品改为 "post_ids": ["123", "456"]。已有 data.posts[].url 可直接交给 media_hunter_download_url 的 url,无需再搜索。
ID 必须来自同一站点;输入最多 50 项,重复 ID 按输入顺序去重,整批格式先校验再联网。去重后一个 ID 下载整部作品,不限文件数;多个 ID 共用 50 个文件目标的预算,复用和失败目标也计入预算。 超过剩余预算的作品整部跳过;模型可根据用户目标、页数及 skipped 原因决定拆分调用,大画廊单独用 post_ids: [该 ID] 下载。不同站点、表里站或原图设置需分次调用。这些决策规则也写在 MCP 工具及参数说明中。
无论单个还是多个,data.requested_ids 都是去重后的 ID 列表,data.downloaded 都是逐作品结果数组(每项含 id、post、files 等,也可能部分失败),errors 和 skipped 分别记录失败、未处理的 ID 及原因。下载不接受 query、rating、limit 或 page,这些条件只在搜索时使用。按链接下载仍直接返回 data.files 等单作品结果。
各工具通过参数 Schema 暴露说明、默认值和范围;下载结果的 files[].path 与 sidecar_path 均为运行服务器的本地路径。操作失败时 success=false / MCP isError=true;部分下载失败仍保留 data 中的文件清单。登录或额度错误先处理原因,其他失败可按原 ID 重试补齐,无需重新搜索。
site 使用 e621、rule34、ehentai 或 pixiv;ExHentai 同样使用 ehentai,可通过每次调用的 use_exhentai 参数切换。
搜索 page 从 1 开始。limit 上限:e621 320、rule34 1000、E-Hentai 100、Pixiv 30。E-Hentai 使用实际的 Next 游标顺序翻页,最多 100 页;深页查询比第一页慢。返回的是站点当前页中符合条件的结果,过滤后可能少于 limit。
rating 语义:e621 为 s/q/e 或完整名称;rule34 为 safe/questionable/explicit;Pixiv 为 all(不限)、safe、r18、r18g,后面三种精确匹配;E-Hentai 为画廊分类,例如 Manga、Non-H。min_score 对 Pixiv 表示收藏数,对 E-Hentai 表示星级。
下载工具的 timeout 是覆盖排队、作品详情、图片页解析、传输与合成的总秒数。省略表示不限制总时长;网络操作仍使用 [network].timeout。
E 站:选择表站、里站和原图
搜索、详情、下载和检查工具均支持 use_exhentai:true 使用里站 ExHentai,false 使用表站 E-Hentai;省略时按配置决定,配置缺省及模板默认均为里站。media_hunter_download_url 是例外:省略时遵循链接域名,也可以显式覆盖。选择里站需要账号 Cookie,不会在失败时悄悄改用表站。旧配置中明确设置的 use_exhentai = false 仍然有效。
两个下载工具还支持 original:默认 false 下载页面图,设为 true 使用页面提供的原图入口。allow_original = true 只表示允许这一选择,不会让每次调用自动下载原图;设为 false 时原图请求会在联网前被拒绝。这两个工具参数仅适用于 E 站。
例如,模型可调用 media_hunter_search 搜索表站:
{"site": "ehentai", "query": "landscape", "rating": "Non-H", "limit": 5, "use_exhentai": false}调用 media_hunter_download 从里站下载指定画廊原图(将 gid/token 换成实际 ID):
{"site": "ehentai", "post_ids": ["gid/token"], "use_exhentai": true, "original": true, "timeout": 300}原图下载可能消耗 FIQ(原图额度)或 GP,具体由账号权益、画廊时间和站点规则决定,参见 E-Hentai 官方下载说明。登录或额度错误会停止后续下载;原图入口失败时不会自动退回缩放图。没有独立原图入口、且页面未标注缩放的图片,使用页面直接提供的源图。
页面图维持原目录;原图保存到该画廊的 original/ 子目录,并保存独立清单。两种模式分别校验和复用,普通图不会被当作已完成的原图。结果中的 post.extra.use_exhentai 和 post.extra.original 表示本次选择。
下载结果
文件流式写入随机
.part临时文件,完成且长度检查通过后原子替换目标文件。失败或取消会清理本次未完成文件。作品下的 JSON sidecar 保存元数据、文件清单、SHA-256 和失败页;ugoira 的帧顺序和延时也会保存。
E-Hentai 每本画廊有独立目录,即使指定相同
subdir也不会覆盖另一本的页码文件。E-Hentai 图片页在每张实际下载前解析;单页失败会记录并继续其他页。认证或配额错误会停止当前作品未开始的下载,也会停止批量任务的后续作品。
部分失败时返回
success: false、error.type: partial_download,同时在data返回已完成文件及错误。按 ID 下载无论单个还是多个,还返回skipped和stop_reason。超时后已完成文件保留;取消中的作品清单写入 sidecar,按 ID 下载结果的
total_files统计已返回的作品结果,不包含取消中作品的残余完成文件。默认在当前输出目录中校验清单的作品身份、文件来源、大小和 SHA-256,复用校验通过的文件,只补下载缺失或损坏的文件。E-Hentai 复用已有页时也会跳过该页的图片地址解析。
MCP 下载工具设置
overwrite=true,或命令行加--overwrite,可以强制重新下载。现有文件仍只在新下载成功后被替换。没有哈希的旧版清单会重新下载一次,建立新版校验记录。files包括新下载和复用文件,文件项的reused表示是否复用;new_files、reused_files分别计数。按 ID 下载的total_files包括两者;attempted_files是进入处理流程的目标数,包含复用及失败目标,仅在去重后多个 ID 时受 50 文件预算约束。取消或重试失败会保留先前清单的文件索引;索引不替代校验,下次复用仍要检查实际文件。续传以完整文件为单位,不保留中断文件的部分字节;跨日期目录、作者目录变更及不同
subdir之间不自动查重。ugoira 的 MP4 输出保留毫秒级帧时长;GIF 按格式限制四舍五入到 10 毫秒、最短 10 毫秒。播放器对短 GIF 帧的显示可能另有限制;精确时序优先使用 MP4 或原始 ZIP。
目录布局:e621/rule34 按日期归档;Pixiv 按作者归档;E-Hentai 按画廊 ID 和标题归档。subdir 是清理后的分组名,不是任意路径。
命令行
uv run media-hunter check
uv run media-hunter search e621 "landscape" --rating safe --limit 5
uv run media-hunter search pixiv "風景" --rating safe --limit 5
uv run media-hunter get pixiv "作品ID"
uv run --extra animation media-hunter download-url "作品页面URL" --timeout 300
uv run media-hunter download-posts e621 "作品ID1" "作品ID2" --timeout 180
uv run --extra animation media-hunter download pixiv "作品ID" --overwrite
uv run media-hunter search ehentai "landscape" --no-use-exhentai --limit 5
uv run media-hunter download ehentai "gid/token" --use-exhentai --original --timeout 300将示例中的 作品ID 替换为实际数字 ID,将 作品页面URL 替换为完整页面链接;E-Hentai 的 ID 使用 "gid/token" 格式。
命令行保留 download 单作品快捷命令;download-posts 接受一个或多个 ID,文件预算及列表返回格式与 MCP media_hunter_download 一致。
指定配置:uv run media-hunter --config "配置文件路径" check。业务结果输出 JSON;操作失败、部分完成或任一站点检查失败时退出码为 1,启动/配置失败为 2。--help 和命令行参数解析错误输出普通文本;Ctrl+C 中断的退出码为 130。配置校验错误会指出字段,例如 network.timeout,不回显该字段中的凭证或其他值。
验证与维护
uv run --extra animation pytest -q
uv run ruff check src tests
uv run ruff format --check src tests测试默认不访问外部站点,覆盖 OAuth 续期、站点解析、下载中断、搜索与下载隔离、明确 ID 批量处理、共享重试预算、取消清理、哈希复用、站点队列、文件上限、镜像画廊分页、表里站并发隔离、原图重定向与额度错误、原图独立复用、真实 ffmpeg 逐帧时序、新旧版 stdio MCP 子进程及 Streamable HTTP 请求。手动联网检查:
uv run --extra animation python tests/live_smoke.py联网检查会下载 e621/Pixiv 的 safe 小样、尝试 E-Hentai Non-H 小样;rule34 只检查搜索、详情与 CDN HEAD。报告和小样写入 .validation/,不写入日常下载根目录。
代码结构:server.py 处理 MCP 接入与协议结果,cli.py 提供命令行;service.py 编排业务;sites/ 负责站点协议;network.py 管理请求;downloader.py 管理校验复用、落盘与合成;progress.py 隔离每次调用的进度回调。新增站点时实现 SiteAdapter 并注册到 MediaService。
许可证
本项目采用 MIT License。
Available Tools
6 toolsdownload_postA
按 ID 下载整部作品,返回文件清单及 sidecar。subdir 为下载根目录下的分组名。 timeout 为覆盖等待、元数据、解析、下载及合成的总秒数。失败文件不会成为最终文件。
| Name | Required | Description | Default |
|---|---|---|---|
| site | Yes | ||
| subdir | No | ||
| post_id | Yes | ||
| timeout | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does substantive work: it defines timeout as covering multiple phases, clarifies subdir semantics, and states that failed files will not become final files. It does not mention auth, overwrite behavior, or partial-file cleanup, but it gives meaningful operational detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action and returning value, followed by terse clarifications of the non-obvious parameters. Every sentence earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is largely complete for a download tool: it covers return contents, parameter semantics for the unusual fields, timeout scope, and failure behavior. The main gaps are undocumented site values and no explicit routing to sibling tools, but an output schema exists and the core call contract is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning for subdir and timeout, and post_id is implied by 'by ID', but the required site parameter is left entirely unexplained, leaving an agent to guess at valid values or the role of site in the request.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'download entire work by ID' and names the return value (file list and sidecar). This distinguishes it from siblings like download_url (by URL) and download_search (by search) without requiring the agent to open schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you have a post ID and want a full download, but it never explicitly says when to prefer this over download_url or download_search, nor gives exclusions. Guidance is implicit rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_searchA
搜索并批量下载,最多尝试 50 个文件。超额画廊整本跳过,可用 download_post 单独下载。 返回 downloaded/errors/skipped;timeout 覆盖搜索和下载全流程。部分完成时 success=false,成功文件仍保留。
| Name | Required | Description | Default |
|---|---|---|---|
| site | Yes | ||
| limit | No | ||
| query | Yes | ||
| rating | No | ||
| subdir | No | ||
| timeout | No | ||
| min_score | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and covers the key runtime facts: the 50-file attempt cap, whole-gallery skip, timeout scope, return fields, and success=false on partial completion while retaining successful files. It does not mention file-location/overwrite behavior or authentication side effects, so it is not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the operation and cap, then give the alternative, return shape, and failure semantics. Every sentence carries unique information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The high-level behavior and edge cases are well covered, and an output schema exists so return values need not be restated. However, with 7 parameters and zero schema descriptions, the lack of guidance on query/site formatting, limit semantics, subdir, rating, and min_score leaves meaningful gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only adds meaning for timeout (covering search and download) and loosely implies a 50-file ceiling. Required site/query plus rating, subdir, limit, and min_score are left unexplained, so the agent must guess their formats or allowed values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase '搜索并批量下载' names a specific action and resource, and the 50-file cap gives clear scope. It explicitly contrasts with 'download_post' for over-limit galleries, so an agent can distinguish it from its siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states that over-limit galleries are skipped and should be downloaded individually via download_post, which is an explicit alternative for a concrete condition. It does not discuss when to prefer plain search or download_url, but the batch-search-and-download intent is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_urlC
识别四个站点的作品页面 URL 并下载。timeout 为总秒数;不接受任意文件直链。
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| subdir | No | ||
| timeout | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It tells the caller that timeout is total seconds and that direct links are rejected, but it does not disclose whether the download creates files, requires authentication, has rate limits, or what side effects occur. This is a significant gap for a download tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the main purpose. However, it omits essential parameter semantics and behavioral details, making it under-specified rather than efficiently complete. It earns a 4 for brevity and structure, not for completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (per context signals), which may compensate for return-value details. But given the tool's complexity (3 params, 0% schema coverage, no annotations), the description is inadequate. It lacks guidance on when to use it, what the four sites are, and how subdir behaves. Completeness is partial at best.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameter meanings. It only clarifies 'timeout' (total seconds). The 'url' parameter's expected format is vaguely implied by 'work page URLs from four sites,' but 'subdir' is completely unexplained. Without schema descriptions, this leaves the agent guessing on two of three parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: identifying and downloading work-page URLs from four specific sites. It distinguishes from generic download tools by specifying the accepted URL type, but it does not name the four sites or explicitly contrast with sibling tools like download_post or download_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a constraint (does not accept arbitrary direct file links) but offers no guidance on when to use this tool versus alternatives like search, get_post, or download_search. It doesn't state prerequisites or exclusions beyond the direct-link restriction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_postC
作品完整元数据。post_id 为数字;ehentai 为 gid/token。
| Name | Required | Description | Default |
|---|---|---|---|
| site | Yes | ||
| post_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It only mentions the outcome (complete metadata) and a post_id format note, with no statement on read-only behavior, side effects, permissions, rate limits, or error conditions. This is insufficient for an agent to safely predict the tool's effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very compact: two sentences. The first sentence delivers the core purpose, the second adds essential post_id semantics. There is no filler, and the structure is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the output schema exists and the parameter count is low, the description leaves critical context unanswered: possible 'site' values, how site and post_id interact, and when to use this tool over the siblings. An agent cannot fully infer the correct invocation just from this description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema coverage at 0%, the description partially compensates by clarifying that post_id should be numeric and that for ehentai it should be a gid/token pair. This adds real meaning to one parameter, but the 'site' parameter is left completely unexplained (allowed values, enums, or relationship to ehentai are not addressed).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states that the tool retrieves complete metadata (完整元数据) for a work, giving a clear verb-resource pairing. It implies a 'get' operation with metadata as the deliverable, but it does not explicitly name how it differs from siblings like download_post or search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus existing alternatives such as search, download_post, or download_url. It lacks explicit conditions, contextual scenarios, or exclusion criteria that would help an agent choose this tool correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
搜索 e621/rule34/ehentai/pixiv。query 使用站点原生标签或关键词。 page 从 1 开始;limit 上限分别为 320/1000/100/30。 rating:e621 s/q/e;rule34 safe/questionable/explicit; pixiv all(不限)/safe/r18/r18g(精确等级);ehentai 为画廊分类。 min_score 为站点分数(Pixiv 收藏数、E-Hentai 星级)。过滤后可能少于 limit。
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| site | Yes | ||
| limit | No | ||
| query | Yes | ||
| rating | No | ||
| min_score | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses page semantics, per-site limits, rating value mappings, min_score meaning, and that filtering can return fewer results than limit. It does not mention auth, rate limits, or read-only status, but it is still more transparent than most.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, scannable, and every clause adds a distinct parameter or per-site fact. It front-loads the action and then efficiently enumerates constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the multi-site complexity and absence of annotations, the description covers the main invocation requirements: per-site ratings, limits, page semantics, and min_score. An output schema exists so return values do not need explanation; the remaining gaps are explicit site parameter values and a concrete query example.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description is the only source of parameter meaning. It explains query, page, limit caps, rating values per site, and min_score. However, accepted site values are only implied by the list 'e621/rule34/ehentai/pixiv' rather than explicitly tied to the site parameter, and no query example is given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: searching e621/rule34/ehentai/pixiv. It clearly frames this as a multi-site search tool and distinguishes it from siblings like get_post and download_post by focusing on query, rating, and limit behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete guidance on how to construct searches: native tags or keywords, 1-based page numbering, per-site limit caps, and per-site rating vocabulary. It does not explicitly name alternatives or state when not to use it, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_checkA
并行检查四站凭证及 API 连通性,每站最多 45 秒;不下载媒体、不返回凭证。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description takes on the transparency burden and does a strong job: it states the parallel behavior, per-call timeout, and explicitly rules out downloading media and returning credentials. It does not mention whether the check has side effects, such as logging or rate-limit consumption, but the read-only intent is reasonably clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the main action, then appends two useful exclusions and a timeout constraint. It is efficient and adds no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that the tool takes no parameters and an output schema exists, the description covers the key operational context: it is parallel, bounded by a 45-second timeout, and intentionally avoids downloading or returning credentials. It does not explain what the four stations are or how results are reported, but these are likely covered by the output schema and domain conventions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the schema already fully covers the input side. The description adds no parameter detail, but none is needed; the baseline of 4 reflects that there is no parameter ambiguity to resolve.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('并行检查' / check in parallel) and a concrete resource ('四站凭证及 API 连通性'), and further clarifies what the tool does not do ('不下载媒体、不返回凭证'). This distinguishes it from retrieval/download-oriented siblings like search, get_post, and download_url.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies this is a preflight/diagnostic tool for checking credentials and connectivity before other operations, but it does not explicitly say when to use it instead of siblings or exclude alternative workflows. The 'does not download media' hint helps somewhat, but there is no direct usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.2.0- First observed
download_post - First observed
download_search - First observed
download_url - First observed
get_post - First observed
search - First observed
self_check
TDQS
Scored across 6 tools
Each tool targets a distinct operation: searching, fetching metadata, downloading by ID, downloading via search, downloading via URL, and running connectivity checks. There is minimal overlap, and the descriptions clarify the boundaries between similar actions.
Tool names follow a clear verb-based snake_case pattern (search, get_post, download_post, download_search, download_url). The slight exception is self_check, which breaks the verb-first pattern but remains readable and unambiguous.
Six tools fit the site-specific media downloader role well, providing all necessary operations without bloat. Each tool serves a clear purpose, and the count is well within the typical and manageable range.
The tool surface covers the full workflow: search, metadata retrieval, single-post downloads, batch search downloads, URL-based downloads, and credential/connectivity self-checks. Download limitations are handled through documented fallbacks, leaving no obvious dead ends for agents.
Maintenance
Related MCP Connectors
Create images and videos from prompts, with options for image mixing, reference images, and start/…
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables browsing, searching, and downloading Pixiv content including illustrations, rankings, recommendations, and user bookmarks. Supports intelligent batch downloads with automatic Ugoira-to-GIF conversion and OAuth 2.0 authentication.4MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to generate images using NovelAI, supporting text-to-image, image-to-image, and tag suggestions.3MIT
- FlicenseAqualityDmaintenanceEnables video metadata extraction and downloading from platforms like TikTok and YouTube via yt-dlp, including direct MP4 links and local downloads with MP3 conversion.3-
- FlicenseNot gradedqualityBmaintenanceEnables downloading videos and audio from various platforms via yt-dlp, with configurable quality, batch downloads, progress tracking, cancellation, and retry mechanisms.-