periscope-mcp
Aggregates content from Bilibili, including Bilibili CC subtitles, as a data source for the news radar and evidence corpus.
Supports Doubao (ByteDance) as an AI provider for scoring, summarization, enrichment, and claim analysis.
Aggregates content from Discourse forums as a data source for the news radar and evidence corpus.
Aggregates content from GitHub as a data source for the news radar and evidence corpus.
Supports Google Gemini as an AI provider for scoring, summarization, enrichment, and claim analysis.
Fetches Google News in adaptive research widening when enabled, providing additional sources for evidence gathering.
Supports Ollama as an AI provider for scoring, summarization, enrichment, and claim analysis.
Supports OpenAI (GPT) as an AI provider for scoring, summarization, enrichment, and claim analysis.
Aggregates content from Reddit as a data source for the news radar and evidence corpus.
Aggregates RSS feeds as a data source for the news radar and evidence corpus.
Aggregates content from Telegram as a data source for the news radar and evidence corpus.
Aggregates content from V2EX as a data source for the news radar and evidence corpus.
Delivers bilingual daily digests via WeChat.
Aggregates content from YouTube as a data source for the news radar and evidence corpus.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@periscope-mcpsearch my evidence corpus for AI safety claims and count independent sources"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
📡 你专属的 AI 新闻雷达:聚合多源信息,自动筛选、去重、富化,生成中英双语每日简报,并把每一天的知识沉淀成可检索、可核查、可研究的证据库。
📋 配置指南 · 🧩 Profile 定制 · 🔎 取证检索机制 · 📊 评测
目录
Related MCP server: Infobroker
简介
好的内容散落在订阅源、论坛与时间线里,而你的注意力是有限的。Periscope 会替你收集、筛选、去重,再把真正值得一读的内容连同背景与社区讨论一起送到你面前。
你的品味决定了你读什么,也决定了你希望从中得到什么。一篇新闻报道需要回答「为什么重要」,一篇工程深度长文需要回答「我能用上什么」。Periscope 的 Profile(画像) 为每一类内容定义各自的评分标准与输出形式,让简报读起来像是为你手工挑选的。
Periscope 回答的不只是「今天有什么值得读」,还有**「这个说法站不站得住」。它在采集与简报之上做三件事:证据语料库(corpus.db 跨运行、跨入口累积,不会次日即忘)、声明级核查(独立信源计数 + 引用反解核验)与多轮共创研究**(长会话里逐轮取证、改范围、深化、定稿),并给出 Web 面板与 MCP(27 个工具)两个入口;再往下是三项支撑机制——证据分层、条目可信度与自适应取证,它们决定了前三项在噪声里是否真的站得住。详见独有层次。
能力总览
能力 | 落在哪 |
多源聚合(14 类源)、Profile 评分、双语日报、邮件 / Webhook / 微信投递、配置向导 |
|
Bilibili / V2EX / Discourse / YouTube 四个源,含 B 站 CC 字幕层 |
|
证据语料库:SQLite + FTS5 + 手写 SimHash 聚簇,跨运行累积 |
|
用户导出入库通路(取不到的源不靠抓取) |
|
证据分层:scraper 声明的类型化 |
|
条目可信度与独立性:可拆解的 trust 分数、noisy-OR 聚合、两道门(跑在 |
|
多轮共创:逐轮动词 + 带 revision/locked/stale 的草稿工件 + 向用户索取输入 |
|
源注册表与抓取基础设施:per-host 令牌桶、可注入时钟、env/cookie 鉴权与过期检测 |
|
取证检索:查询扩展 + 向量路 + RRF 融合 |
|
声明级核查 + 评级一致率工具 |
|
长会话研究:子问题树、崩溃续跑、缺证时自适应加宽 |
|
报告骨架(背景调查 / 市场调研 / 方法探索)+ 引用反解核验 |
|
检索消融评测(Recall / Precision / nDCG / MRR) |
|
Web 面板 |
|
本项目的立场一句话:把来源放宽到论坛与流媒体,信息质量必然下降;全部工作在于让"降下去的质量"被分层、检索、计数与引用核验重新补回来,并用消融表说明补回了多少。
核心特性
下面每条都是本项目自带的能力:
📡 聚合你的信息源 — RSS、Hacker News、Reddit、Telegram、X、GitHub、财经资讯等 10 类,再加 Bilibili、V2EX、Discourse、YouTube 四类与 B 站 CC 字幕层,共 14 类。
🎯 判断什么值得读 — 用 Profile 定义评分细则,为每个 Profile 设置阈值,并在同一 Profile 内合并重复报道。
🧩 为每篇内容定制处理方式 — 通过 Markdown 提示词与 JSON block 定义,自由组合摘要、背景、解决方案或要点。
💬 不止于标题 — 在有助于解释事件时,自动补充联网检索到的背景与社区讨论。
⚖️ 兼顾你的所有兴趣 — 限制简报总长度与各分类占比,避免某一热门话题挤占其余内容。
📬 在你习惯的地方阅读 — 生成中英双语 Markdown 简报,发布到 Pages,或通过邮件、Webhook、微信投递。
🧠 让每一天都可积累 — 所有采集过的内容进入证据语料库,可全文检索、做声明核查,并作为长会话研究的素材;证据分层保证评论与回复只作线索,不充当信源。
🧯 诚实的降级 — 没有 LLM key 也能运行:语料照常增长、证据照常确定性关联,报告会如实标注「尚未回答」而不是编造内容。
🔍 找不到就换打法 — 一个子问题不再只有一次查询:放宽词条、换来源族、让模型改写问法,每步写入
research_actions,报告尾部列出「取证尝试」,把"语料里没有"和"问法不对"分开。✅ 引用可自查 — 成品报告可被反向核验:幽灵引用、指向不存在条目、只靠人群发言支撑的引用,都会被点名(
src/corpus/citations.py)。
一份简报,多种读法
Profile 是一套可复用的编辑规则:什么内容该收录、什么值得保留、该写成什么样。 内置示例如下:
你关注的内容 | Profile | 你将得到 |
科技新闻 |
| 事件与背景,必要时附影响分析与社区讨论 |
工程深度长文 |
| 背景、解决方案与可落地的要点 |
AI 创作者素材 |
| 摘要,必要时附热点切入角度与内容思路 |
财经资讯 |
| 公司动态与市场脉络 |
你可以为某个信息源指定 Profile,也可以让 AI 自动选择。想要不一样的风格?通常无需改 Python,直接改编现有 Profile 即可。定制你的 Profile →
截图
工作原理
Profile 决定处理方式,运行期配置反映你的阅读偏好。 每个条目只会路由到一个 Profile。分析、过滤与去重在富化之前完成,随后入选条目按 Profile 分组形成简报。结构与图元(通路、会话状态机、一轮时序)见 docs/architecture.md——那三张图与代码是双向对齐的,画进去的名字由测试核对。
一次完整运行(periscope --hours 24)的阶段顺序如下:
采集 — 并发抓取所有已启用数据源,每个源独立成败,结果同时写入证据语料库并重算近似重复聚簇。
跨源去重 — 规范化 URL(剥离
utm_*、gclid等跟踪参数),合并指向同一内容的多源条目。分析 — 分类路由到 Profile,由 AI 完成评分、理由、摘要与标签。
选择与过滤 — Profile 阈值过滤 → AI 主题去重 → 均衡配额(
digest配置)。富化 — 第二遍 AI,按 Profile 定义的 block 生成多语言产物,可调用联网检索与历史检索工具。
声明核查 — 将高分条目蒸馏为原子声明,关联语料证据并评级(尽力而为)。
摘要与投递 — 程序化渲染多语言 Markdown 日报,保存到
data/summaries/,并按配置发布或投递。
你可以通过 CLI 运行完整流水线,也可以让 AI 助手经由 MCP 调用其中的各个阶段。
独有层次
只回答「今天有什么值得读」、并且到了第二天就遗忘的工具,产出的是一天的信息,不是可复查的证据。Periscope 在采集与简报之上做了八项增强,让"来源放宽必然掉下去的质量"能被补回来:
证据语料库(
corpus.db) — 所有曾经采集过的条目都会带 SimHash 指纹与近似重复聚簇被持久化存储,可用 SQLite FTS5(对 CJK 友好)检索。知识会跨运行累积,而不是在生成摘要后就蒸发。声明级正确性核查 — 高分条目会被蒸馏为原子化、可核查的声明;每条声明都与语料证据关联并评级(supported / contested / unsupported + 置信度),评级每轮有预算上限。
contested不再只是一个标签:claim_contradictions会记下是哪两条声明、靠哪些条目互相冲突。长会话研究 — 提出一个问题,会得到一棵分解后的子问题树,逐题对照语料取证,并输出带引用的 Markdown 报告。追问会在同一会话上跨天、跨重启迭代(状态保存在 SQLite 中)。
多轮共创(逐轮动词) — 研究循环的顶层动词是
AskUser | Rescope | Deepen | Finalize。一轮step只重算被点名的分支,返回这一轮改了什么;报告是带 revision 的草稿工件,用户改过或锁定的章节在后续重算中只被标记为陈旧、不会被覆盖;缺输入时系统会主动提问并把会话停在awaiting_user(没有可用 LLM 时回退只deepen/finalize,不会连环追问)。处处诚实降级 — 没有 LLM key?语料照常增长,证据照常确定性关联,报告会写明*「尚未回答」*并附上已收集的证据,而不是编造内容。默认 LLM 是免费的 Agnes 层(
agnes-2.5-flash),且每次调用都经过持久化响应缓存与限流,因此崩溃恢复运行与重复提示词都不消耗额外额度。证据分层(类型化字段,不靠字符串约定) — 抓取器把评论、楼层回复、视频字幕声明为
ContentItem.sections:primary(作者亲写,可承载声明)与community(人群发言,只能当线索),每段带author/provenance/locator(可引用到具体一层)。声明只从前者蒸馏、证据只与前者关联、报告引用的也是前者。旧的标记反解只作为tiering="marker"的消融档与老库回填存在,src/scrapers/里已不允许出现任何分层标记字面量。可用analysis.claimable_only/research.claimable_only关闭,用于对照。条目可信度与独立性 — 每条入库内容算一个可拆解的信任分:
σ(源先验, 作者等级, 交叉支持, 可核验实体, 新鲜度, 来源方式 − 模板度),特征与分数一起落库,"为什么这条被当成证据"随时答得出来。独立信源先按簇折叠重复内容、再数不同的(source_type, publisher),解析不出发布者的条目不投票;声明级聚合用 noisy-OR 而不是求和(求和没有上界,够多的低质源能把任何结论刷成 supported)。阈值θ是手工先验、分诊门默认关闭,等人工标注到位才谈校准。自适应取证加宽 — 一个子问题不再只有一次查询:先放宽词条,再换没查过的来源族,然后让模型改写查询,最后(显式开启时)用 GDELT / Google News 现采一轮。每次尝试都写进
research_actions,未回答的子问题会在报告里列出「取证尝试:动作(+新增条数)」,把"语料里确实没有"和"这次的问法没查到"区分开。
这些能力由 CLI、Web 面板与 MCP 三个入口共享同一份 corpus.db,因此会话与证据跨入口、跨重启都可见。各层的实际收益见检索与取证评测。
驱动它
# 终端:一次采集(fetch -> score -> digest -> corpus + claims)
uv run periscope --hours 24
# Web 面板:证据库 / 研究报告 / 核查台(http://localhost:8790)
uv run periscope-web --data-dir data
# MCP:面向任意 MCP 客户端的 27 个工具(ps_research_start、ps_research_step、ps_corpus_search 等)
uv run periscope-mcp
# 取不到的源:用户自己导出,按声明的层级入库(不联网、不碰验证码与签名)
# 样例负载见 data/export.example.json(把内容换成你账号本来就能看见的东西)
uv run python scripts/import_corpus.py --file data/export.example.json --data-dir data --dry-run
uv run python scripts/import_corpus.py --file export.json --data-dir data配置键:corpus、analysis、research(见 data/config.example.json)。把 .env 指向 AGNES_API_KEY 即可获得完整体验;不配置任何 key,其余功能也能运行。
快速开始
1. 安装
方式 A:本地安装
git clone https://atomgit.com/NLY22/periscope.git
cd periscope
# 使用 uv 安装(推荐)
uv sync
# 需要测试/开发依赖时
uv sync --extra dev
# 或使用 pip
pip install -e .Periscope 需要 Python 3.11 及以上(见 pyproject.toml 的 requires-python)。
dev 目前是 pyproject.toml 中的可选 extra,因此安装 pytest 等开发依赖请使用 uv sync --extra dev。
如果你想启用可选的 OpenBB 财经新闻源,还需要安装它的 extra:
uv sync --extra openbb如果 openbb 在你的机器上拉取到没有 wheel 的包,可只用二进制方式手动安装 SDK:
uv pip install --only-binary=:all: openbb openbb-benzinga方式 B:Docker
git clone https://atomgit.com/NLY22/periscope.git
cd periscope
# 可选:首次运行前用逗号分隔的 extras 构建
docker compose build --build-arg EXTRAS=openbb periscope-collect基于 trafilatura 的正文抽取已包含在基础安装中。twitter extra 还需要 Playwright 浏览器与系统依赖,当前 Dockerfile 并未安装。
2. 配置
方式 A:交互式向导(推荐)
uv run periscope-wizard向导会询问你的兴趣(例如「LLM 推理」「嵌入式」「Web 安全」),并自动生成 data/config.json。CLI 选项见交互式向导。
方式 B:手动配置
cp .env.example .env # 填入你的 API Key
cp data/config.example.json data/config.json # 定制你的信息源最小手动配置示例:
{
"ai": {
"provider": "openai",
"model": "gpt-4",
"api_key_env": "OPENAI_API_KEY"
},
"sources": {
"rss": [
{
"name": "Simon Willison",
"url": "https://simonwillison.net/atom/everything/",
"profile": "tech-news"
}
]
},
"processing": {
"profiles_dir": "profiles",
"default_profile": "tech-news",
"profile_settings": {
"tech-news": {
"threshold": 7.0,
"topic_dedup": true
}
}
}
}信息源显式指定 profile 时会直接使用该 Profile;省略或设为 "auto" 时,由 AI 将条目与所有可用 Profile 进行匹配;设为数组(如 ["tech-news", "finance-news"])则把 AI 匹配限定在这些 Profile 内。Profile 的结构与行为见处理画像。诸如评分阈值、主题去重之类的单 Profile 用户偏好应写在 processing.profile_settings,而不是 Profile 文件里。
均衡简报(可选)
限制最终简报的规模,避免某一分类独占结果。分类来自信息源配置,例如 sources.rss[].category。
{
"digest": {
"max_items": 20,
"category_groups": {
"ai": {
"limit": 5,
"categories": ["ai-news", "ai-tools", "machine-learning"]
},
"finance": {
"limit": 5,
"categories": ["finance", "business", "equities"]
}
},
"default_group": "other",
"default_group_limit": 3
}
}分组限额在 Profile 过滤之后、富化之前生效。省略 category_groups 与 max_items 时,不施加均衡简报限制。
api_key_env 必须是环境变量的名字,而不是 API Key 本身。把真正的密钥放进 .env:
OPENAI_API_KEY=sk-your-key使用 Gemini 时,改为 GOOGLE_API_KEY:
{
"ai": {
"provider": "gemini",
"model": "gemini-2.0-flash",
"api_key_env": "GOOGLE_API_KEY"
}
}data/config.json 中任意字符串值都可以用 ${VAR_NAME} 引用环境变量,适用于 ai.base_url、私有 RSS 源地址、Webhook 端点或自定义请求头模板等。
完整参考见配置指南。
3. 运行
A. 本地安装
uv run periscope [OPTIONS]B. Docker
# 采集器,运行一次
docker compose run --rm periscope-collect --hours 24
# Web 面板,端口 8790
docker compose up -d periscope-web选项 | 默认值 | 说明 |
| 取 | 抓取最近 N 小时的内容 |
|
| 数据目录路径 |
|
| 配置文件路径 |
|
| 日志级别(DEBUG/INFO/WARNING/ERROR/CRITICAL) |
--data-dir 会改变状态目录,包括摘要、订阅者以及默认配置位置;--config 只改变配置文件。生成的报告保存在 data/summaries/(若设置了 --data-dir 则为 <data-dir>/summaries/)。两者组合使用及自定义配置位置的初始化方式见配置路径。
4. 自动化(可选)
Periscope 适合用系统定时器调度,例如 cron 或 systemd timer;Docker Compose 也可直接配合定时任务使用(docker compose run --rm periscope-collect --hours 24 配 cron 即可)。
仓库里另有 .github/workflows/tests.yml。本平台不执行任何 GitHub 语法的 workflow(已实测:PR 的 check_tasks_num 为 0),那条文件只对在 GitHub 上跑的人有意义;在这个平台上要定时运行,就用任何能执行 shell 的调度器,按下面的命令排程即可。
支持的 AI 提供商
默认使用免费的 Agnes 层,其余提供商可任选其一,或通过 provider_chain 组合成降级链。
提供商 | 默认模型 | API Key 环境变量 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 无需 Key |
设置 ai.provider_chain(逗号分隔的提供商列表)即可启用链式回退:遇到 429/限流、401/403/额度不足、502/503 或空响应时,自动切换到下一个提供商。
支持的数据源
数据源 | 抓取内容 | 评论 / 附加能力 |
Hacker News | 按分数排序的热门帖子 | 是(前 N 条评论) |
RSS / Atom | 任意 RSS 或 Atom 源 | 可选全文抽取 |
子版块 + 用户帖子 | 是(前 N 条评论) | |
Telegram | 公开频道消息 | — |
Twitter / X | 用户时间线 + 关键词搜索(Apify);另支持 Playwright + Cookie 的免 token 模式 | 是(前 N 条回复) |
GitHub | 用户事件与仓库 Release | — |
OpenBB | 按自选股/提供商抓取公司财经新闻(可选 SEC filings) | — |
OSS Insight | 开源趋势仓库 | — |
GDELT | 匹配搜索词的全球新闻 | — |
Google News | 经 RSS 的新闻搜索 | — |
Bilibili | 热门视频(标题、统计、热门评论) | 是;有 CC 字幕时提取转录 |
V2EX | 节点主题 + 热门主题 | 是(含回复) |
Discourse | 任意 Discourse 论坛的最新主题 | 是(含楼层讨论) |
YouTube | 经官方公开 atom feed 抓取频道更新 | 标题 + 描述文本 |
投递渠道
Periscope 可以通过多种方式发布或投递生成的简报:
渠道 | 作用 |
邮件订阅 | 向订阅者发送每日简报,并通过 SMTP/IMAP 处理订阅/退订请求 |
Webhook 通知 | 把成功或失败结果推送到飞书/Lark、钉钉、Slack、Discord 或任意自定义 Webhook 端点 |
微信通知 | 通过 iLink Bot 在扫码登录并收到你的消息后发送简报;受微信回复条数限制 |
进阶用法
Web 面板
uv run periscope-web --data-dir data # 默认 http://localhost:8790三栏式界面:证据库(全文检索)、研究报告(长会话研究)、核查台(按 verdict 展示声明与证据)。研究报告栏按轮次组织:推一轮 按钮、每节的「已锁定」与「新版已变 · 未覆盖」标记、待回答请求卡片;报告正文会列出未评级声明的原因(低于分诊门 / 证据不足 / 无可核验发布者 / 可信度不足)。顶部可一键触发采集。API 文档位于 /api/docs。
MCP 服务
uv run periscope-mcp以 stdio 方式运行,提供 27 个 ps_* 工具与 7 个 periscope:// 资源(4 个固定 + 3 个带参数的模板,实测握手),覆盖配置校验、分阶段流水线、运行产物、证据检索、声明核查、研究会话与 Webhook 通知。详见 MCP 工具说明 与客户端配置。
微信与 Webhook 命令行
uv run periscope-wechat login # 扫码登录 iLink Bot
uv run periscope-wechat status # 查看登录状态
uv run periscope-wechat test # 测试发送(可 --dry-run 预览)
uv run periscope-webhook --dry-run # 预览 Webhook 请求配置参考
完整示例见 data/config.example.json,密钥见 .env.example。主要配置块:
配置块 | 作用 |
| 提供商、模型、 |
| 各数据源及其分类、Profile 指定 |
| 默认时间窗口 |
| 简报总长度与分类配额( |
|
|
| 图标风格 |
| 正文抽取器配置 |
| 邮件订阅与 SMTP/IMAP 设置 |
| Webhook 端点、平台适配、消息模板与投递语言 |
| 微信投递开关、语言与分块大小 |
| 证据语料库: |
| 声明核查: |
| 长会话研究: |
| 取证检索: |
项目结构
periscope/
├── src/
│ ├── main.py # CLI 主入口
│ ├── orchestrator.py # 流水线编排与阶段复用
│ ├── models.py # 数据模型与配置定义
│ ├── scrapers/ # 14 类数据源抓取器
│ ├── ai/ # 多 provider 客户端、分析/富化/摘要/本地化
│ ├── processing/ # Profile 加载、历史检索、内容处理
│ ├── corpus/ # 证据语料库(SQLite + FTS5 + SimHash)
│ ├── analysis/ # 声明级核查
│ ├── research/ # 长会话研究
│ ├── web/ # Web 面板(FastAPI + 单文件三栏 UI)
│ ├── mcp/ # MCP 服务与运行产物存储
│ ├── services/ # 邮件 / Webhook / 微信投递
│ ├── setup/ # 交互式配置向导
│ ├── extractors/ # 正文抽取(trafilatura)
│ └── storage/ # 配置与摘要存储
├── profiles/ # 内置 Profile(tech-news / tech-blog / ai-creator / finance-news)
├── docs/ # 文档与站点资源
├── tests/ # 测试
├── data/ # 配置示例与运行数据
├── scripts/ # 辅助脚本
├── Dockerfile
├── docker-compose.yml
└── pyproject.toml开发与测试
uv sync --extra dev # 安装开发依赖
uv run pytest # 运行全部测试
uv run pytest tests/test_corpus.py # 运行单个测试文件
uv run python scripts/eval_retrieval.py # 检索消融表(docs/evaluation.md)
uv run python scripts/eval_retrieval.py --tiering marker # 复现分层前的 A 档
uv run python scripts/eval_multiturn.py # 多轮调用数比值 / 轮次 / 灌水曲线(不联网)关于 CI:这个平台上没有自动执行。 .github/workflows/tests.yml 是 GitHub 语法的配置,AtomGit 不执行它(每个 PR 的 check_tasks_num 都是 0,已实测确认)。所以:
那条文件保留着,迁到 GitHub 或支持该语法的平台就生效;内容仍然是可信的验收脚本(Linux + Windows 各跑一遍全量,再跑一次检索 harness,只看能否复现,不在 CI 里断言指标数值)。
但在当前平台,验证是提交者的责任:本地跑
uv run pytest与相关 harness,把实测数字写进 PR 正文。本项目不接受「有 CI 兜底」作为质量证据,PR 模板性的「测试通过」需要能复现的命令。
测试位于 tests/,覆盖流水线各阶段、各数据源抓取器、证据语料库与声明核查、研究会话、Web 面板与 MCP 服务等。
安全与可靠性
SSRF 防护 — 抓取与正文抽取前校验 URL 的 scheme、主机与端口,解析后要求所有地址为公网,并逐跳校验重定向。
路径逃逸防护 — 摘要与运行产物写入前校验目标路径位于允许目录内。
凭据脱敏 — MCP 返回的有效配置会递归脱敏
token、secret、password等字段,Webhook 日志同样脱敏。降级而非崩溃 — LLM 缓存读写失败、单个数据源抓取失败、声明评级失败都不会中断整条流水线。
文档
指南 | 说明 |
AI 提供商、信息源、Profile、过滤、邮件、Webhook、微信与 MCP 配置 | |
Profile 路由、提示词、运行期过滤偏好、富化 block 与工具 | |
Periscope 如何评估与排序新闻条目 | |
各数据源抓取器细节与扩展说明 | |
RSS 源的全文抽取 | |
消融表、标注口径、多轮成本与灌水曲线、已知不足 | |
五条命令自己把效果重测一遍:每步该看到什么、哪些"怪现象"是正常的、真模型腿取不到时怎么降级 | |
已合入 | |
噪声从哪来、类型化分层怎么声明、可信度与独立性怎么重算、找不到时怎么加宽、多轮草稿怎么保住了用户的字、怎么复现 | |
三张结构性图(spec §8 的 1–3 号):广源→可用证据的通路、研究会话状态机( | |
免费方案:Playwright + 自己账号的 cookie 抓推文(Apify 订阅的替代路径) | |
面向 MCP 兼容客户端的 27 个工具参考 | |
为什么这么改(spec v4,§15 是交付记录)、 |
项目状态
日报闭环(本项目自带):多源采集、Profile 驱动的分析与富化、去重、评论摘要、双语生成、邮件 / Webhook / 微信投递、Docker 部署、MCP 集成与配置向导。
在此基础上,证据语料库、声明级核查与多轮共创研究是三项核心能力,证据分层、条目可信度与自适应取证是三项支撑机制,Web 面板是独立入口(见独有层次)。
后续计划(与 docs/superpowers/specs/ 里设计 spec 的 §15.5「还剩什么」一致):
声明 verdict 的人工标注 50–100 条 → 报 macro-F1 与按独立信源数分桶的一致率,再用
roc_thresholds()校准 θ_s / θ_triage。这是本项目唯一"能力已实现、效果未主张"的一块:工具已就位(scripts/eval_claims.py --export/--score),缺的是标注本身,在此之前阈值是手工先验、分诊门默认关闭源可达性探针:判别逻辑已经做成可复跑的工具 ——
src/sources/reachability.py+scripts/spike_sources.py(判定分为pass / list_only / blocked_captcha / signed_required / blocked_auth / error,离线用httpx.MockTransport测,不加--online不发任何请求)。贴吧 S2 的第 1 步结论(列表可达、楼层不可达)现在是测试里可复现的判定而不是散文;小红书 S1 的第 1 步同样不需要账号,一条命令就能跑并把结果落到data/eval/reachability_results.json:
uv run python scripts/spike_sources.py --source xiaohongshu --url <公开笔记页 URL> --online
uv run python scripts/spike_sources.py --source tieba --kw <吧名或关键词> --online(--online 是唯一会让它真的发请求的开关;去掉它就是打印判定计划。)第 2–3 步才需要登录态或浏览器。S1 不通过时的降级路径已经实现(用户导出 → scripts/import_corpus.py / ps_corpus_import / POST /api/import,分层靠声明),所以"取不到的源"不再是死路,只是不自动。
P3 图文 → 文本通路(OCR / VLM):条件执行,卡在 S1 结论;VLM 描述只能当线索,不进声明蒸馏
支持更多数据源类型,例如 Discord
平台侧未决:GitHub 语法的 workflow 在本平台不被执行(护栏与测试要人工过一遍)、仓库 issue 开关只能在网页打开
在 AtomGit 上发布 Release;发布到 PyPI,支持
pip install
贡献
欢迎贡献:issue 与 PR 请开在 https://atomgit.com/NLY22/periscope。规范见 CONTRIBUTING.md,行为准则见 CODE_OF_CONDUCT.md,安全问题见 SECURITY.md。动手前建议先读 docs/retrieval.md 与 docs/evaluation.md —— 新增能力要能进消融表,否则只是加功能。信息源与 Profile 的改动直接开 PR 或 issue。
许可证
Available Tools
27 toolsps_corpus_importA
Import a user export ({"items": [...]}) into the evidence corpus.
Each item declares its own authorship tiers through sections, so crowd
text from an export cannot become evidence. Use this for sources Periscope
must not scrape; see src/mcp/README.md for the payload shape.
dry_run=true reports what would be accepted and which items are rejected,
without storing anything.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | ||
| payload | Yes | ||
| tiering | No | sections | |
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does add real context: crowd text cannot become evidence because each item declares its own authored tiers, and dry_run stores nothing. It omits idempotency/overwrite behavior on the target corpus and permission requirements for a write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, followed by the tiering guarantee and dry_run note. Structurally efficient, with only the README redirect detracting slightly from self-containment.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be explained, but the payload is a nested free-form object the description defers to src/mcp/README.md, and two path parameters are unexplained. For a tool with nested input and tiering logic, that redirect leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 5 params, so the description must compensate. It explains payload shape ({"items": [...]}) and dry_run semantics, and references `sections` for tiering, but config_path, periscope_path, and the tiering parameter itself are left undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource ("Import a user export ... into the evidence corpus") and clarifies the intent versus scraping-based sources. It implicitly distinguishes itself from the read-oriented siblings like ps_corpus_search, though it never names an alternative tool explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete when-to-use rule: "Use this for sources Periscope must not scrape," and describes the dry_run preview path as the safe way to test acceptance. No explicit when-not or named-alternative routing is provided, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_corpus_recentC
Most recently published items in the corpus, optionally one source.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| config_path | No | ||
| source_type | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it only discloses that results are ordered by publication recency. It says nothing about pagination, the default limit behavior, auth/config prerequisites, or what the returned items represent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler. It is efficient, though the terseness contributes to the coverage gaps elsewhere rather than being a virtue in isolation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but with zero annotation coverage, four undocumented parameters, and no usage guidance, the description leaves too much for the agent to infer about how to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and four parameters exist, so the description must compensate; it only vaguely gestures at source filtering (likely source_type). The limit, config_path, and periscope_path parameters are undocumented in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase names the resource (items in the corpus), the ordering/scope (most recently published), and an optional narrowing (one source), which implicitly separates it from the search sibling. However, there is no explicit verb and no direct naming of ps_corpus_search as the alternative, so the distinction is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all, and no mention of when to prefer ps_corpus_search or ps_filter_items over this tool. The only hint is 'optionally one source', which gestures at filtering but does not tell the agent when that option is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_corpus_searchC
Full-text search over every item ever collected (BM25, CJK-aware).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the retrieval algorithm (BM25, CJK-aware), which is useful, but says nothing about ranking, default result count, pagination, or whether config_path/periscope_path change what corpus is searched.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, and the core resource is front-loaded. It is efficient, though very slightly under-specified rather than verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but with zero annotation coverage, 0% parameter documentation, and no usage guidance, the definition leaves too much for an agent to infer about scope, limits, and path parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with 4 parameters, so the description must compensate and largely does not. It implies query is a full-text search string, but limit, config_path, and periscope_path receive no explanation in either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Full-text search over every item ever collected') and adds implementation character (BM25, CJK-aware) that tells the agent what matching behavior to expect. The phrase 'every item ever collected' implies scope beyond the sibling ps_corpus_recent, but no sibling is named explicitly, so differentiation remains implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no alternatives named, despite several plausible siblings (ps_corpus_recent, ps_corpus_import, ps_filter_items, ps_fetch_items). An agent must guess whether to search here or pull recent items elsewhere.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_corpus_statsB
Evidence-corpus overview: stored items by source, clusters, runs, plus claim and research-session counters. The 'what do we already know' probe.
| Name | Required | Description | Default |
|---|---|---|---|
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the shape of the output (counts broken down by source, clusters, runs, claims, sessions), which is real value, but it never states that the call is read-only/side-effect free or whether it needs configuration to be valid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the returned-content inventory front-loaded and the usage cue at the end. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the description covers the semantic content of the overview well. However, with two completely undocumented optional path parameters, the definition is not quite complete for an agent deciding how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters (config_path, periscope_path) have 0% schema description coverage, so the description must compensate and it does not — neither parameter is mentioned. An agent has no guidance on what these paths mean or when to supply them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and the content it returns (stored items by source, clusters, runs, plus claim and research-session counters), which clearly distinguishes it as an aggregate-overview tool rather than a retrieval tool. It never names a sibling (e.g. ps_corpus_search) as the alternative, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing phrase 'The "what do we already know" probe' implies this is a situational-orientation call to make before searching or running research. That is useful implied context, but there is no explicit when-to-use versus when-not, and no alternative is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_enrich_itemsC
Enrich filtered items into the enriched stage.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| config_path | No | ||
| source_stage | No | filtered | |
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden and largely fails. It does not state whether this mutates state (implied by advancing to a new stage), whether it requires network/LLM calls, whether re-running is idempotent, or what auth/run preconditions apply. Only the source→target stage movement is communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is structurally clean. But it is sized far below what a 4-parameter pipeline-mutation tool needs, so brevity tips into under-specification rather than efficient conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema relieves the description of explaining return values, but the remaining gaps are large: no annotations, four undocumented parameters, and no description of what enrichment produces or the ordering constraints with ps_filter_items. As a stage-mutating pipeline tool it is significantly under-described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate and does not. The run_id, config_path, and periscope_path parameters are entirely unexplained, and while 'filtered' faintly echoes the source_stage default, the parameter itself is never described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It pairs a verb (enrich) with a resource (items) and names the target stage (enriched), which loosely distinguishes it from siblings like ps_fetch_items, ps_score_items, and ps_filter_items. However, 'enrich' is undefined jargon — the agent gets no sense of what enrichment actually does to the items. Purpose is identifiable but vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites (e.g., that filtering must precede enrichment, or that a run must exist), and no alternatives named. The phrase 'filtered items' implies a sequencing but nothing is explicit. An agent must infer the pipeline position.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_fetch_itemsC
Fetch and deduplicate content into the raw stage.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | ||
| run_id | No | ||
| sources | No | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it discloses almost nothing: it does not say whether fetching is idempotent, how deduplication keys are chosen, whether existing raw items are overwritten, or what permissions/rate limits apply. 'Deduplicate' is the one real behavioral hint, which keeps this above a 1.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is structurally sound. It is under-specified rather than bloated, so the size itself is not a defect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with zero schema coverage and no annotations, the description omits everything an agent needs to call it correctly - source selection, time window semantics, run scoping. The presence of an output schema excuses it from describing return values but not from the input-side gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across five parameters, and the description explains none of them. hours, run_id, sources, config_path, and periscope_path are all opaque from the description; only 'sources' is loosely implied by 'fetch content.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb (fetch) and a resource (content) and scopes it to the 'raw stage,' which hints at its position as the first pipeline stage relative to ps_score_items/ps_enrich_items. However, it never says what content or from where, so it is only loosely distinguishable from sibling ingest-style tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no preconditions, and no mention of alternatives such as ps_corpus_import or ps_run_pipeline. The agent must infer the pipeline position entirely from the phrase 'raw stage.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_filter_itemsC
Filter scored items into the filtered stage.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| threshold | No | ||
| config_path | No | ||
| topic_dedup | No | ||
| source_stage | No | scored | |
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden yet only implies a mutation ('into the filtered stage') without stating whether items are moved, copied, or overwritten, whether the operation is idempotent, or what permissions/run state it needs. No side effects or reversibility are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence is front-loaded and free of padding, which is good. But at this length it is under-specified rather than concise, leaving no room for the behavioral or parameter context a 6-parameter tool needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a pipeline-stage tool with 6 parameters, no annotations, and 0% schema documentation, the description is far too thin. The presence of an output schema excuses it from explaining return values, but nothing about inputs, ordering, or side effects is covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must compensate and largely fails. The words 'scored' and 'filter' faintly gesture at the source_stage and threshold concepts, but config_path, topic_dedup, and periscope_path are entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a specific verb (filter) and resource (items) and names a target stage, so the general action is inferable. However, 'filtered stage' is pipeline jargon with no explanation of what filtering criteria apply, and it does not distinguish itself from pipeline siblings like ps_score_items or ps_enrich_items beyond the stage label.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when this stage should run relative to siblings (e.g., after ps_score_items, before ps_enrich_items), no prerequisites, and no mention of alternatives. The default source_stage='scored' hints at ordering but the description never states it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_generate_summaryC
Generate a markdown summary from a stage.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| language | No | zh | |
| config_path | No | ||
| source_stage | No | ||
| periscope_path | No | ||
| save_to_periscope_data | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It is silent on whether the operation writes anything (the save_to_periscope_data flag implies a possible side effect), what permissions are needed, or whether an existing summary is overwritten — all material for a generation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is not padded, but one terse line is under-specification rather than conciseness for a six-parameter tool with side-effect flags. Nothing is front-loaded because there is essentially only one clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, but the description leaves the rich parameter set, side-effect semantics, and relationship to sibling summary/run tools entirely unaddressed. This is inadequate for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across six parameters, and the description explains none of them. Critical fields like source_stage, config_path, periscope_path, and the language default ('zh') are completely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb ('Generate') and output artifact ('markdown summary') are stated, giving a rough sense of the operation. However, 'from a stage' is ambiguous and the definition does not distinguish this from the sibling ps_get_run_summary, so an agent cannot confidently pick between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer (e.g., ps_get_run_summary). The agent is left to infer usage entirely from the name and the stage parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_get_claimB
One claim in full: verdict, evidence rows, source independence.
| Name | Required | Description | Default |
|---|---|---|---|
| claim_id | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It lists what comes back but says nothing about permissions, behavior on an unknown claim_id, pagination of evidence rows, or side effects — significant gaps for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded phrase with no wasted words. Its brevity borders on under-specification rather than verbosity, but structurally it is clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out (the description repeats them anyway). For a simple get-by-id tool this is near-adequate, but the undocumented config_path/periscope_path parameters and absence of any behavioral notes leave real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters. The description never mentions claim_id, config_path, or periscope_path, so the two optional path parameters are completely undocumented in both schema and description and the agent must guess at their purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('one claim') and the payload it returns (verdict, evidence rows, source independence), so an agent knows this is a single-entity detail fetch. It is distinguishable from ps_list_claims by the word 'One', but siblings are never named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the 'one claim in full' framing suggests calling this after a list/search returns a claim_id. There is no statement of when to prefer it over ps_list_claims or ps_corpus_search, and no prerequisites or failure conditions mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_get_metricsB
Read in-memory server metrics.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It implies a safe read via 'Read' but discloses nothing about authentication, freshness, rate limits, or whether it is truly non-destructive, leaving a significant gap for a server-facing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero waste. It is appropriately sized for a simple read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema means return values need not be explained. However, with no annotations and many sibling tools that also provide statistics, the description is too thin to fully guide an agent—it never says what 'server metrics' covers or how this differs from corpus or run summaries.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 per the rubric. The description adds no parameter information, but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('in-memory server metrics'), making the tool's basic function clear. However, it does not distinguish this tool from sibling stats tools such as ps_corpus_stats or ps_get_run_summary, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of alternatives, and no prerequisites or conditions. The sentence only declares what the tool does, leaving the agent to infer when to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_get_run_metaC
Read run metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden, and it only conveys the bare implication that "Read" is non-destructive. It says nothing about error behavior for an unknown run_id, whether the run must be complete, or what class of data counts as metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three words is technically concise, but this is under-specification rather than efficiency: there is no front-loaded scope, no disambiguation, and no actionable information. Brevity here costs the agent more than it saves.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need not be described, but for a tool sitting among many similarly named run-level getters the definition leaves the key question unanswered: what distinguishes this metadata call from run summary, stage, or metrics. It is incomplete for a 27-sibling toolset.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter run_id is documented nowhere. The description could have said where run_id comes from (e.g., ps_list_runs), but it adds no semantics beyond the parameter's self-explanatory name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a verb and resource ("Read run metadata"), so the operation is identifiable, but "metadata" is undefined and the description never distinguishes it from close siblings like ps_get_run_summary, ps_get_run_stage, or ps_get_metrics. An agent cannot tell from the text which of these four run-inspection tools returns what.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all: no statement of prerequisites, no condition under which this tool is preferred over ps_get_run_summary or ps_list_runs, and no exclusions. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_get_run_stageC
Read items from a run stage.
| Name | Required | Description | Default |
|---|---|---|---|
| stage | Yes | ||
| run_id | Yes | ||
| max_items | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. The word 'Read' implies a non-mutating operation, but nothing states pagination/truncation behavior for max_items (default 200), error behavior for invalid run_id or stage, or what happens when a stage is empty or incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence is efficient and front-loaded, but here brevity shades into under-specification rather than disciplined conciseness — the sentence earns its place but leaves obvious room for useful additional detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but with zero annotations, 0% parameter coverage, and an undefined domain term ('stage'), the definition is not complete enough for an agent to call this confidently among 26 siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only loosely echoes two of the three parameter names ('run', 'stage') without adding format, accepted values, or enumeration of valid stages. max_items is entirely unaddressed, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description has a clear verb ('Read') and resource ('items from a run stage'), but never defines what a 'stage' is or which stages exist, which matters given siblings like ps_get_run_meta and ps_fetch_items that operate on adjacent data. It is a plausible but vague purpose, not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this instead of ps_fetch_items, ps_get_run_meta, or ps_get_run_summary, nor any stated prerequisites (e.g. run must be in a particular state). The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_get_run_summaryC
Read a generated run summary.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| language | No | zh |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden and delivers very little. It implies a read operation but does not disclose whether a summary must be pre-generated, what happens if none exists, language behavior, or any response characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no wasted words, which is structurally sound. But its brevity reflects under-specification rather than efficient communication of necessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists (so return values need not be described), the description omits the prerequisite relationship to ps_generate_summary, the meaning of the language parameter, and any distinction from peer run getters. For a tool operating in a 26-tool suite with several near-identical getters, this is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 2 parameters, and the description adds nothing: neither run_id nor the language parameter (default 'zh', which selects the summary language) is explained. The description does nothing to compensate for the undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read a ... run summary'), so the basic operation is clear. However, it does not distinguish itself from closely related siblings such as ps_get_run_meta, ps_get_run_stage, and ps_get_metrics, all of which are read-only getters on the same run object. The purpose is adequate but not disambiguated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance and never mentions the relationship to ps_generate_summary (must a summary exist first?) or to the other run getters. An agent has no criteria for choosing this tool over its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_list_claimsC
Claims by pipeline status (extracted | linked | graded) with verdicts, confidence, trust, ungraded_reason and independent-source counts.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No | graded | |
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It lists returned fields but says nothing about read-only safety, required permissions, pagination behavior, or side effects, leaving key behavioral traits undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler and is front-loaded around the resource and filter. It could be marginally tighter by omitting return-field details that the output schema already covers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four parameters at 0% schema coverage and no annotations, the description is incomplete. It does not explain most parameters, lacks usage guidance relative to siblings, and omits behavioral context, even though the output schema makes return-value explanation less necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for the status parameter by listing extracted | linked | graded, but provides no explanation for limit, config_path, periscope_path, or the default of graded, leaving three of four parameters semantically undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool lists claims by pipeline status and enumerates the status values, making the core purpose understandable. It does not, however, explicitly differentiate this tool from sibling list/search tools such as ps_get_claim or ps_corpus_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. The status filter implies a usage context, but the description never states when to choose it over ps_get_claim, ps_corpus_search, or other list tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_list_runsC
List recent runs and stage states.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: it does not state ordering, what 'recent' means (time window? top-N?), pagination behavior, or that results are capped by the limit parameter's default of 20. It is plausibly a read-only list, but that is inferred, not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence is front-loaded and contains no filler, which is structurally good. It is terse to the point of under-specification rather than genuinely concise, so it earns only a middle score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value shape need not be explained. But with zero annotations and zero parameter documentation, the description leaves the agent without the scoping, ordering, and limit semantics needed to call this listing tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'limit' parameter, and the description never mentions it, its default of 20, or the maximum. The parameter name is fairly self-explanatory, but the description does nothing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description has a verb ('List') and a resource ('runs'), plus a vague secondary object ('stage states'), so the general purpose is discernible. However, it gives no basis for choosing it over close siblings such as ps_get_run_meta, ps_get_run_stage, or ps_get_run_summary, and 'recent' and 'stage states' are undefined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no conditions, and no mention of alternative tools like ps_get_run_stage or ps_get_run_summary that also expose run information. The reader is left to infer that this is the entry-point listing tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_answerB
Reply to a question the session asked (skip: true to decline), then the
loop continues. A parked session is not stuck: this is how you unpark it.
| Name | Required | Description | Default |
|---|---|---|---|
| skip | No | ||
| answer | No | ||
| request_id | Yes | ||
| session_id | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the control-flow effect (loop resumes, session unparks) and the decline path, which is genuinely beyond the schema. However, it omits auth requirements, idempotency/retry behavior, and what happens if answer is omitted or the request_id is stale.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and the decline flag, with no filler. It is appropriately short for a control-flow call, though the second sentence is more encouragement than specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The core interaction is conveyed, but with 6 parameters at 0% schema coverage and no annotations, the definition leaves key invocation details (what answer/request_id should contain, path params) under-specified for a session-driving tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must compensate and largely does not. It explains only `skip`; `answer`, `request_id`, `session_id`, `config_path` and `periscope_path` carry no meaning beyond their titles, leaving an agent unsure what to pass for the required identifiers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action (reply to a session's question) and its effect (the loop continues), which distinguishes it from generic corpus/pipeline siblings. It does not explicitly contrast itself with near neighbors like ps_research_followup or ps_research_step, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the trigger ('a question the session asked') and gives one decision path (`skip: true` to decline), plus the 'how you unpark it' framing. But it never states when to prefer this over ps_research_followup/ps_research_step or what prerequisites (e.g. a parked session, valid request_id) must hold.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_draftB
Read the research document as a versioned draft: revision, origin
(render | user_edit | merge), and per-section state — locked, stale, and
the evidence item ids behind each section. Omit revision for the latest.
| Name | Required | Description | Default |
|---|---|---|---|
| revision | No | ||
| session_id | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does meaningful work: it discloses the payload structure, the enumeration of origin values (render | user_edit | merge), and the per-section flags (locked, stale) plus evidence linkage. It stops short of error behavior when a revision is missing or any auth/session requirements, so it is strong but not exhaustive for an annotation-free tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact statements, front-loaded with the payload description and ending with the revision default. No filler, and each clause (origin values, section state, evidence ids) adds distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so explaining return values is not strictly required, yet the description covers them well while leaving the non-revision parameters (config_path, periscope_path) unexplained in an environment with 0% schema coverage and no annotations. Adequate for calling the tool with a session id, incomplete for the full parameter surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 4 parameters, so the description must compensate and largely does not. It explains only `revision` ('omit for the latest'); session_id, config_path, and periscope_path get no meaning at all, leaving three of four parameters documented nowhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Read) and resource (the research document as a versioned draft) and enumerates what the read yields: revision, origin, and per-section lock/staleness/evidence state. That is well beyond a restatement of the name. It does not, however, distinguish itself from siblings like ps_research_list or ps_research_status, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage cue is 'Omit `revision` for the latest,' which is a parameter default hint rather than selection guidance. There is no statement of when to prefer this over ps_research_list or ps_research_status, and no preconditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_editB
Rewrite one section as the author. Stored as its own revision and locked, so a later recompute marks it stale instead of overwriting your prose.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| section_id | Yes | ||
| session_id | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: the edit is stored as its own revision, is locked, and a later recompute marks it stale rather than overwriting it. This is genuinely non-obvious and valuable, though it omits any auth/permission or concurrency details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses with the action front-loaded and the persistence/locking caveat second. Nothing is redundant, though the second clause is carrying a lot in one sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the revision/lock semantics are well covered. But for a 5-parameter mutation tool with zero schema descriptions and no annotations, the description leaves parameter meaning and any precondition (session validity, section existence) unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across five parameters, so the schema offers no guidance at all. The description only indirectly hints at section_id ('one section') and body ('your prose') and says nothing about session_id, config_path, or periscope_path, leaving three parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: rewriting a single section as the author, which is distinguishable from the sibling generation tools (ps_research_draft, ps_research_answer). However, it never names an alternative explicitly, so the agent must infer the boundary between editing and drafting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: 'Rewrite one section' suggests it applies to an existing section the author wants to own, but there is no explicit when-to-use condition and no mention of the sibling tools it competes with (ps_research_draft, ps_research_step).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_followupB
Push back on an existing session: the message can narrow, expand or challenge; the sub-question tree is revised and the report re-rendered.
| Name | Required | Description | Default |
|---|---|---|---|
| message | Yes | ||
| session_id | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose real side effects well – the sub-question tree is revised and the report re-rendered – which is more than most. It omits permissions/auth requirements, whether the revision is reversible or overwrites prior state, and the failure mode if the session does not exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that is front-loaded with the core action ('push back on an existing session') before the effect clause. No filler, though it packs three behaviors into one sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the mutation behavior is described. Still, for a no-annotation mutation tool the description leaves parameter meaning and sibling discrimination uncovered, so an agent cannot confidently route to it over ps_research_edit/answer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 4 parameters, so the description must compensate. It adds meaning only for 'message' (narrow/expand/challenge), leaving session_id, config_path, and periscope_path entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('push back on') and resource ('an existing session') and even clarifies the effect: the message can narrow, expand or challenge, and the sub-question tree is revised and report re-rendered. It is clearly a follow-up/refinement tool. However it never names or contrasts with close siblings such as ps_research_edit or ps_research_answer, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Existing session' and 'push back' imply this is only for continuing an already-started research session, which is useful implicit context. But there is no explicit when-to-use guidance, no exclusions, and no routing to the obvious alternatives (ps_research_edit, ps_research_answer, ps_research_step) that share this niche.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_listC
Recent research sessions with their questions and statuses.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: 'recent' hints at default ordering but there is no mention of how results are limited, whether config_path/periscope_path are required for discovery, or pagination behavior. The output schema covers the return shape, which relieves some burden, but scoping and prerequisite behavior remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler; it is front-loaded and wastes no words. Its brevity is appropriate structurally, though it is under-specified rather than genuinely concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, but for a discovery tool with three undocumented parameters and no annotations, the description leaves the agent without the scoping, ordering, or configuration context needed to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters (limit, config_path, periscope_path) have zero schema description coverage. The description adds no meaning to any of them — in particular it never explains what config_path or periscope_path are for, which is the most ambiguous part of the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase names the resource (research sessions) and what it returns (questions, statuses), giving an implied list operation. However it is a noun fragment with no verb, and it does nothing to distinguish itself from siblings like ps_research_status or ps_corpus_recent, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus ps_research_status, ps_research_start, or any other sibling. The agent must infer from the name alone that this is the enumeration entry point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_startB
Open a long-session research task: decompose the question, investigate each sub-question against the corpus, return a cited markdown report. State persists — follow up later, even after restarts.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses that this is a long-session task with persistent state and a markdown report return, which matters for an agent. However, it says nothing about cost, latency, permissions, or whether the task is asynchronous vs blocking — significant gaps for a long-running operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and the pipeline it triggers. No filler; the persistence note earns its place. Slightly dense but well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format needn't be explained, and the description does cover the task lifecycle. But with zero annotation coverage, 0% parameter documentation, and no mention of the two config-path arguments, an agent lacks enough to invoke this long-running tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 3 parameters. The description never mentions config_path or periscope_path, leaving two parameters entirely undocumented in both schema and prose. "question" is inferable from the workflow narrative, but the two config paths are opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Open a long-session research task") and enumerates the internal workflow (decompose, investigate sub-questions, return a cited markdown report). It is distinguishable from read-only siblings like ps_research_status or ps_research_list, though it never names an alternative to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase "follow up later, even after restarts" implies this is the entry point that pairs with ps_research_followup, but no explicit when-to-use or when-not-to-use guidance is given. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_statusC
Session state: sub-question tree with statuses/answers/evidence ids, turn history, and the current rendered report.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It describes the returned session-state contents, but does not explicitly state that the operation is read-only, whether it has side effects, or what permissions are required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the resource and lists the relevant state fields. It is concise, though its terseness leaves several important details unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is less critical, but the description still omits usage guidance and all parameter semantics for a 3-parameter tool. For an agent selecting among many research siblings, it is not complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has three parameters and 0% schema description coverage, so the description must compensate. It never mentions session_id, config_path, or periscope_path, adding no meaning beyond the bare parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource and enumerates the session-state contents (sub-question tree, statuses/answers/evidence ids, turn history, rendered report), so an agent can tell this retrieves research-session status. However, it does not explicitly differentiate this tool from siblings such as ps_research_list or ps_research_step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no when-to-use guidance or alternatives. It does not say whether to call this before/after ps_research_step, or how it relates to ps_research_list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_research_stepB
Advance a research session by ONE round and get back the delta.
The loop chooses a verb each round: deepen a branch, fold the message into
the question tree, finalise — or ask you for something only you can supply
(then pending_request is non-null and the session parks). Returns the
changed section ids and the new draft revision, not a whole re-rendered
tree, so a round costs only what it touched.
| Name | Required | Description | Default |
|---|---|---|---|
| message | No | ||
| session_id | Yes | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose useful traits: one round per call, verb selection, the parking/pending_request behavior, and the delta-only return that keeps cost proportional to changes. It omits auth requirements, error handling, idempotency/re-entrancy semantics, and what happens if the session is already finalised.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and return shape are front-loaded in the first sentence, and the following sentences add the verb set and payload semantics without much waste. It is slightly dense with parenthetical asides but remains appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the description does convey the round/parking model. For a 4-parameter tool with zero schema description coverage and no annotations, however, the gaps in parameter meaning and usage routing leave it short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all four params, yet it explains none of them. It implicitly references 'the message' (fold the message into the question tree) and mentions pending_request (which is not even a parameter), but session_id, config_path, and periscope_path are entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('advance a research session by ONE round') and clarifies the scope is a single step returning a delta rather than a full tree. It is clear what the tool does, but it never explicitly names or contrasts the sibling tools (ps_research_followup, ps_research_answer, ps_research_draft) that an agent might confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the loop semantics: the tool advances a round and may park with a non-null pending_request when input is needed. However, there is no explicit statement of when to call this versus followup/answer/status, nor any prerequisite guidance, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_run_pipelineC
Run fetch -> score -> filter -> enrich -> summarize in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | ||
| enrich | No | ||
| sources | No | ||
| languages | No | ||
| threshold | No | ||
| config_path | No | ||
| topic_dedup | No | ||
| periscope_path | No | ||
| save_to_periscope_data | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the entire behavioral burden, and it discloses only the internal stage ordering. It says nothing about side effects, despite a save_to_periscope_data parameter defaulting to false that implies an optional write, nor about runtime, external calls, idempotency, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and free of waste, but it is under-specified rather than concise for a nine-parameter pipeline tool with no annotations. Its brevity leaves the agent guessing on nearly every invocation detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a nine-parameter orchestrator with zero parameter documentation and no annotations the definition is materially incomplete. Side effects and defaults around persistence (save_to_periscope_data) go entirely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Nine parameters with 0% schema description coverage receive no explanation whatsoever in the description. Only 'enrich' is loosely echoed by the stage list, and nothing clarifies hours, threshold, sources, languages, config_path, topic_dedup, periscope_path, or save_to_periscope_data, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb ('Run') and enumerates the composite stages (fetch, score, filter, enrich, summarize), which map cleanly onto sibling tools like ps_fetch_items, ps_score_items, ps_filter_items, ps_enrich_items, and ps_generate_summary. The phrase 'in one call' implicitly distinguishes it from running those stages individually, though it never says so outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or named alternative; the only cue is 'in one call,' which implies this is the batched shortcut for the individual stage tools. An agent can infer the intent but must guess at prerequisites, cost, or when the single-stage tools are preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_score_itemsD
Score a stage into the scored stage.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| config_path | No | ||
| source_stage | No | raw | |
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and carries almost none of it. 'Into the scored stage' hints at a state transition/mutation of a run, but there is no disclosure of permissions, reversibility, side effects, or idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loaded, but the brevity is under-specification rather than conciseness. Every word is a restatement of the tool name and contributes no actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though an output schema exists (removing the need to document return values), the definition provides no annotations, no parameter semantics, and no usage context for a four-parameter pipeline tool. An agent cannot reliably decide when or how to invoke it from this description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four parameters (run_id, config_path, source_stage, periscope_path), and the description says nothing about any of them. The defaults 'raw' and null behavior are equally unexplained, so nothing compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description restates the tool name with a circular phrase: 'Score a stage into the scored stage.' It never says what scoring means, what gets scored, or what items (ps_score_items) have to do with stages. It names a verb and a noun but no meaningful resource semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no reference to any alternative. The mention of 'stage' faintly implies membership in a pipeline workflow that includes siblings like ps_run_pipeline or ps_get_run_stage, but nothing tells the agent when to pick this tool over them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_send_webhookB
Send a webhook notification with the given variables.
Uses the webhook URL (from environment variable), request_body template, and headers from the Periscope config. Template variables #{date}, #{language}, #{important_items}, #{all_items}, #{result}, #{timestamp}, #{summary} are replaced in the URL and request_body before sending.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| result | No | success | |
| summary | No | ||
| language | No | zh | |
| all_items | No | ||
| config_path | No | ||
| periscope_path | No | ||
| important_items | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose important behavior: the webhook URL comes from an environment variable, headers/body template come from the Periscope config, and template variables are substituted in both URL and body before sending. However, it omits failure behavior, what happens if the URL or config is missing, auth requirements, and whether the call blocks or retries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences that front-load the action and then the mechanics; the parameter list is dense but earns its place given 0% schema coverage. Minimal filler, though the second sentence packs a lot into one clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. But for a side-effecting external notification with no annotations, the description should cover failure modes and the unresolved config_path/periscope_path parameters; it does enough to call the tool but not enough to predict its behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the schema provides no parameter meaning, and the description compensates by naming most substitution variables (#{date}, #{language}, #{important_items}, #{all_items}, #{result}, #{summary}) and explaining that they are replaced in URL/body. It leaves config_path and periscope_path completely unexplained (and #{timestamp} has no matching parameter), so it only partially covers 8 undocumented params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination ("Send a webhook notification") and clarifies that it uses a configured URL, request_body template, and headers. It is unambiguous about what the tool does, though it does not differentiate itself from any sibling (no sibling handles webhooks, so the ambiguity risk is low).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives, no prerequisites stated as conditions, and no exclusions. The mention of config/env-var sourcing implies a setup requirement but never tells the agent under what circumstances to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ps_validate_configC
Validate Periscope config and required environment variables.
| Name | Required | Description | Default |
|---|---|---|---|
| sources | No | ||
| check_env | No | ||
| config_path | No | ||
| periscope_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It states what is checked but not how failure is surfaced (exception vs. error list), whether anything is mutated on disk, or whether env-var checks can be skipped; the only behavior hint, check_env, lives in the schema rather than the prose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It is appropriately sized, though its brevity is partly under-specification rather than economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but with zero annotation coverage and 0% parameter documentation the description leaves an agent without enough context to invoke the four optional arguments meaningfully or to anticipate the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for four parameters. The phrase 'config' loosely hints at config_path and 'environment variables' at check_env, but 'sources' and 'periscope_path' are entirely unexplained in either place, so the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Validate Periscope config') and expands scope to 'required environment variables'. It is unambiguous what the tool does, though it makes no attempt to relate itself to any sibling (none of the siblings overlap, so differentiation is less critical here).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to run this tool, e.g. before ps_run_pipeline or after changing config, and no prerequisites or alternatives are named. The reader must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
27 tool updates
v0.1.0- First observed
ps_corpus_import - First observed
ps_corpus_recent - First observed
ps_corpus_search - First observed
ps_corpus_stats - First observed
ps_enrich_items - First observed
ps_fetch_items - First observed
ps_filter_items - First observed
ps_generate_summary - First observed
ps_get_claim - First observed
ps_get_metrics - First observed
ps_get_run_meta - First observed
ps_get_run_stage - First observed
ps_get_run_summary - First observed
ps_list_claims - First observed
ps_list_runs - First observed
ps_research_answer - First observed
ps_research_draft - First observed
ps_research_edit - First observed
ps_research_followup - First observed
ps_research_list - First observed
ps_research_start - First observed
ps_research_status - First observed
ps_research_step - First observed
ps_run_pipeline - First observed
ps_score_items - First observed
ps_send_webhook - First observed
ps_validate_config
TDQS
Scored across 27 tools
Most tools target a clear resource+action: individual pipeline stages (fetch/score/filter/enrich), distinct run reads (meta/stage/summary), and distinct research lifecycle steps. Minor overlaps exist between ps_generate_summary vs ps_get_run_summary, ps_corpus_stats vs ps_get_metrics, and ps_corpus_search vs ps_corpus_recent, but descriptions differentiate them adequately.
All tools share a consistent ps_ snake_case prefix, which makes the set highly scannable. The only deviation is that some names are verb-first (ps_get_claim, ps_list_runs, ps_validate_config) while others are domain-first (ps_corpus_search, ps_research_start), but the pattern remains predictable within each domain.
At 27 tools this is heavier than typical, though the server genuinely spans several domains (corpus, claims, pipeline stages, research sessions, webhook). The eight individual pipeline-stage tools plus ps_run_pipeline create some redundancy, so the surface could be trimmed.
Coverage is broad and lifecycled: corpus ingest/search/stats, claim inspection, full pipeline plus per-stage control, research start/followup/step/draft/edit/answer/status/list, and webhook notify. Only minor gaps exist (e.g., no explicit corpus/claim deletion or run cancellation), which agents can work around.
Maintenance
Related MCP Connectors
MCP-native web evidence and claim verification: cited, source-grounded evidence for AI agents.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Source-traced evidence research for AI agents. We organise the evidence; you decide.
Viral-content intelligence for AI agents — 7 read-only MCP tools, evidence-layer scoring.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides local-first web intelligence over MCP with tools for search, fetch, crawl, extract, cache, find-similar, research, and autonomous agent loops, requiring no API keys.101,063 npm5,446AGPL 3.0
- AlicenseAqualityAmaintenanceAn MCP-native research assistant — search backend, intent-routing, and writing pipeline that carries a question through search, verification, and writing, delivering whatever research you need: quick answers, verified sources, or a finished document.47388 npmMIT
- AlicenseAqualityBmaintenanceEnables AI clients to perform autonomous multi-agent deep research through MCP tools, including deep research, quick search, and retrieval of archived Markdown reports with live web search and source citations.4MIT

genpark-deep-researchofficial
FlicenseNot gradedqualityBmaintenanceEnables AI agents to perform deep research across multiple sources, arbitrating consensus while evaluating source authority and detecting contradictions. It integrates via MCP with clients like Claude Desktop and Cursor to deliver deterministic, structured research dossiers.7-