cqupt-notice
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cqupt-noticeFetch the latest CQUPT academic affairs notices and summarize the key points."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CQUPT 教务处通知爬取与 AI 推送系统
基于 DrissionPage + MCP + AstrBot 的重庆邮电大学教务处通知自动爬取、AI 分析与 QQ 推送系统。
每天早上 8 点,自动抓取教务处最新通知,由 AI 分析重要性并生成友好文案推送到你的 QQ。
📋 目录
Related MCP server: Generic MCP Gateway
系统架构
┌─────────────────────────────────────────────────────────────┐
│ AstrBot (QQ 机器人) │
│ │
│ ┌──────────┐ 调用工具 ┌──────────────────────┐ │
│ │ FutureTask│ ──────────────▶│ MCP Client (内置) │ │
│ │ (定时任务) │ └──────────┬───────────┘ │
│ └──────────┘ │ │
│ ▲ │ HTTP/STDIO │
│ │ 推送 ▼ │
│ ┌────┴─────┐ ┌──────────────────┐ │
│ │ AI Agent │ ◀── 通知数据 ──── │ MCP Server │ │
│ │ (大模型) │ │ (get_latest_notices) │ │
│ └────┬─────┘ └────────┬─────────┘ │
│ │ 生成文案 │ 爬取 │
│ ▼ ▼ │
│ ┌─────────┐ ┌──────────────────┐ │
│ │ QQ 推送 │ │ DrissionPage │ │
│ └─────────┘ │ + Chrome │ │
│ └────────┬─────────┘ │
└─────────────────────────────────────────┼───────────────────┘
▼
https://jw.cqupt.edu.cn/tzgg.htm工作流程
定时触发:AstrBot 的 FutureTask 每天 08:00 唤醒 AI Agent
调用工具:Agent 调用 MCP 工具
get_latest_notices爬取解析:MCP Server 用 DrissionPage 绕过 WAF 爬取通知,解析后去重
AI 分析:大模型分析每条通知的重要性、比赛建议、关键信息
QQ 推送:生成友好文案推送到 QQ
技术选型与原理
为什么用 DrissionPage 而不是 requests?
目标页面 jw.cqupt.edu.cn 使用了 加速乐 WAF,首次访问会返回 JS 挑战页面(HTTP 412),
要求浏览器执行 JS 计算后才能获得真实内容。
方案 | 结果 | 原因 |
| ❌ 失败 | 无法执行 JS,只能拿到挑战页 |
| ❌ 失败 | 新版加速乐防护已升级,旧绕过手段失效 |
| ❌ 失败 | 使用 CDP 协议,自动化特征明显,被 WAF 识别 |
DrissionPage + Chrome | ✅ 成功 | 通过启动参数隐藏自动化特征 |
DrissionPage 反检测原理
启动 Chrome 时传入两个关键参数:
co.set_argument("--disable-blink-features=AutomationControlled") # 移除 navigator.webdriver 标识
co.set_argument("--headless=new") # 新版无头模式,更接近真实浏览器这样 WAF 的 JS 检测脚本执行时,navigator.webdriver 返回 undefined(而非 true),
浏览器指纹看起来像真实用户,从而通过挑战、获得有效 Cookie。
为什么用 MCP?
MCP(Model Context Protocol)是 AstrBot 官方推荐的外部工具扩展方式。
它把"爬取通知"这个能力封装成标准化工具 get_latest_notices,让 AI Agent 可以像调用函数一样调用它。
环境要求
依赖 | 版本要求 | 说明 |
Python | >= 3.10 | |
Chrome / Chromium | >= 100 | DrissionPage 会自动调用系统 Chrome |
AstrBot | 最新版 | QQ 机器人框架,需支持 MCP |
大模型 API | 任意 | 如 河图、OpenAI、通义千问等 |
💡 Chrome 只需正常安装即可,DrissionPage 会自动查找。 如果自动查找失败,可在
config.json中指定chrome_path。
快速开始
第一步:安装依赖
# 1. 克隆本项目
git clone https://github.com/Habapure/cqupt-notice-pusher.git
cd cqupt-notice-pusher
# 2. 创建虚拟环境(推荐)
python -m venv venv
# Windows
venv\Scripts\activate
# Linux / macOS
source venv/bin/activate
# 3. 安装依赖
pip install -r requirements.txt第二步:测试爬虫
在接入 AstrBot 之前,先独立测试爬虫是否正常工作:
# 基本测试(爬取当天通知,不标记已推送)
python tests/test_crawler.py
# 查看页面上所有通知(不按日期过滤)
python tests/test_crawler.py --all
# 有头模式(能看到浏览器操作,便于调试)
python tests/test_crawler.py --no-head
# 爬取最近 3 天的通知
python tests/test_crawler.py --days 3✅ 预期输出:
============================================================
CQUPT 教务处通知爬虫 - 独立测试
============================================================
目标 URL : https://jw.cqupt.edu.cn/tzgg.htm
无头模式 : True
日期范围 : 最近 1 天
============================================================
[1/4] 正在爬取页面...
✅ 爬取成功,HTML 长度: 12345
[2/4] 正在解析通知列表...
✅ 解析到 20 条通知
[3/4] 日期过滤(最近 1 天)...
✅ 过滤后剩余 3 条通知
[4/4] 去重检查...
已推送记录: 0 条
✅ 去重后剩余 3 条新通知
============================================================
爬取结果
============================================================
1. [2026-09-21] 🆕 新
标题: 关于XXX的通知
链接: https://jw.cqupt.edu.cn/info/1012/69051.htm
...⚠️ 如果爬虫失败,请先参考 常见问题 排查。
第三步:配置 MCP Server
复制配置文件模板:
cp config.example.json config.json编辑
config.json(通常无需修改,默认即可):
{
"target_url": "https://jw.cqupt.edu.cn/tzgg.htm",
"days_to_fetch": 1,
"headless": true,
"chrome_path": null,
"record_file": "pushed_records.json",
"page_load_timeout": 30
}测试 MCP Server 能否正常启动:
# 测试爬取功能(不启动 MCP 服务)
python mcp_server.py --no-mark
# 测试 MCP Server 能否启动(stdio 模式,会阻塞,Ctrl+C 退出)
python mcp_server.py mcpMCP Server 支持两种传输模式:
模式 | 命令 | 适用场景 |
stdio(默认) |
| AstrBot 和爬虫在同一台机器 |
streamable-http |
| AstrBot 和爬虫在不同机器 |
第四步:接入 AstrBot
📖 完整图文指南见
astrbot/future_task_guide.md,以下是快速版。
方式 A:stdio 传输(同机部署,推荐)
确保 AstrBot 已安装并能正常运行
在 AstrBot 管理后台 → MCP 管理 → 添加 MCP Server:
字段
值
名称
cqupt-notice传输方式
stdio启动命令
python(或虚拟环境 python 的绝对路径)命令参数
["/你的绝对路径/mcp_server.py", "mcp"]⚠️ 路径必须是绝对路径
Windows:
E:\projects\cqupt-notice-pusher\mcp_server.pyLinux:
/home/user/cqupt-notice-pusher/mcp_server.py
如果用了虚拟环境,启动命令写虚拟环境的 python:
Windows:
E:\projects\cqupt-notice-pusher\venv\Scripts\python.exeLinux:
/home/user/cqupt-notice-pusher/venv/bin/python
保存并连接
方式 B:streamable-http 传输(跨机器部署)
在 MCP Server 所在机器启动:
python mcp_server.py mcp --transport streamable-http --host 0.0.0.0 --port 8000在 AstrBot 管理后台添加 MCP Server:
字段
值
名称
cqupt-notice传输方式
streamable-httpURL
http://MCP服务器IP:8000/mcp保存并连接
验证连接成功
在 AstrBot 日志中看到以下内容即表示成功:
✅ MCP 服务器 cqupt-notice 连接成功
已注册工具: get_latest_notices
已注册工具: get_notice_count第五步:设置定时推送
在 AstrBot 管理后台进入「主动任务」/「FutureTask」页面
新建任务:
字段
值
任务名称
重邮教务处通知早报执行时间
每天
08:00投递目标
你的 QQ(私聊或群聊)
系统提示词
见下方
系统提示词(复制
astrbot/system_prompt.md的内容):你是一个「重庆邮电大学教务处通知推送助手」。你的职责是每天定时获取教务处最新通知, 分析每条通知的重要性,并以友好、简洁的格式推送给用户。 工作流程: 1. 调用 MCP 工具 get_latest_notices 获取当天的最新通知列表 2. 对每条通知分析:重要程度(高/中/低)、比赛建议、关键信息 3. 按固定格式生成推送文案 推送文案格式: 🌅 早安!今天是 X 月 X 日,以下是教务处最新通知: 📋 通知1:《通知标题》 重要程度:⭐⭐⭐ 比赛建议:值得参加 / 不建议参加 截止日期:XXXX-XX-XX 🔗 原文链接:https://... —— 重邮教务处通知早报保存并启用任务
测试:点击「立即执行」,你的 QQ 应收到推送消息 🎉
配置说明
config.json 各字段说明:
字段 | 类型 | 默认值 | 说明 |
| string |
| 通知公告页 URL |
| int |
| 抓取最近 N 天的通知(1 = 仅当天) |
| bool |
| 是否无头模式运行 Chrome |
| string/null |
| Chrome 可执行文件路径,null 为自动查找 |
| string |
| 已推送记录文件路径 |
| int |
| 页面加载超时(秒) |
去重机制
为避免重复推送同一条通知,系统维护一份已推送记录文件 pushed_records.json:
每次爬取后,对比通知 URL 是否已推送过
只返回未推送过的通知
get_latest_notices被调用后,自动将返回的通知标记为已推送
重置去重记录
如果需要重新推送所有通知,删除 pushed_records.json 即可:
rm pushed_records.json # Linux / macOS
del pushed_records.json # Windows常见问题
Q: 爬虫返回空列表 / 爬取失败
可能原因 & 解决方案:
WAF 拦截
将
config.json中headless设为false,观察浏览器是否被拦截确认 Chrome 版本 >= 100
页面结构变化
教务处改版后 HTML 结构可能变化
用
python tests/test_crawler.py --no-head打开浏览器检查页面根据实际 HTML 调整
mcp_server.py中的parse_notices函数
网络问题
确认服务器能正常访问
jw.cqupt.edu.cncurl https://jw.cqupt.edu.cn/tzgg.htm测试连通性
Q: MCP 工具在 AstrBot 中找不到
先在命令行运行
python mcp_server.py确认脚本无报错检查 AstrBot MCP 配置中的路径是否为绝对路径
检查 AstrBot 日志中是否有 MCP 连接错误
确认
mcp包已安装:pip show mcp
Q: 推送内容为空
这是正常现象。如果当天没有新通知,或所有通知都已推送过,AI 会回复 "今天教务处没有新通知哦~"。
可删除 pushed_records.json 后重新测试。
Q: DrissionPage 找不到 Chrome
在 config.json 中手动指定 Chrome 路径:
{
"chrome_path": "C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe"
}常见 Chrome 路径:
Windows:
C:\Program Files\Google\Chrome\Application\chrome.exemacOS:
/Applications/Google Chrome.app/Contents/MacOS/Google ChromeLinux:
/usr/bin/google-chrome或/usr/bin/chromium-browser
项目结构
cqupt-notice-pusher/
├── README.md # 本文件,完整教程
├── mcp_server.py # MCP Server 主程序(爬虫 + 解析 + 去重 + 工具)
├── requirements.txt # Python 依赖
├── config.example.json # 配置模板(复制为 config.json 使用)
├── .gitignore
├── astrbot/
│ ├── system_prompt.md # AstrBot AI Agent 系统提示词
│ └── future_task_guide.md # FutureTask 详细配置指南
└── tests/
└── test_crawler.py # 独立爬虫测试脚本许可证
MIT License
致谢
DrissionPage - 强大的 Python 浏览器自动化库
AstrBot - 多平台 QQ 机器人框架
MCP - Model Context Protocol
本项目仅供学习交流使用,请遵守学校网站的使用条款,不要对目标网站造成过大压力。
This server cannot be deployed
Maintenance
Related MCP Connectors
Give an AI agent eyes on the web: turn any feed, page, or stream into deduplicated change events.
AI Agent Source Registry. 288K+ curated sources for agentic search and discovery.
AI agent gateway with web fetching, data extraction, crypto pricing, and x402 payments
Search for AI agents. Closes the LLM-cutoff gap: CVEs, papers, frontier AI, prediction markets.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to discover, fetch, and manage RSS feeds from 1500+ websites via RSSHub, with subscription groups and content filtering.3MIT
- FlicenseNot gradedqualityDmaintenanceA general-purpose MCP gateway integrating tool management, memory, reminders, and multi-channel messaging (Telegram/QQ), enabling autonomous AI agents.-
- AlicenseNot gradedqualityCmaintenanceAggregates security research articles from WeChat, QAX Butian, and Xianzhi community, with optional KimiCode web search. Supports smart deduplication, article fetching, and anti-bot measures.MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to interact with Moodle LMS, fetching assignments, grades, deadlines, course content, and syncing to Obsidian, with WhatsApp alerts and class-slot filtering.-