wechat-article-fetcher
by mrlong0129
README.md
# wechat-article-fetcher
把**公开的**微信公众号文章(`mp.weixin.qq.com`)读成干净的 Markdown 或 JSON 的命令行工具 / Python 库。
A small CLI and Python library that turns public WeChat Official Account articles into clean Markdown or JSON.
作者 Yomin Ma · 博客 [yomin.love](https://yomin.love) · GitHub [@mrlong0129](https://github.com/mrlong0129) · MIT License
```bash
wxfetch https://mp.weixin.qq.com/s/<文章ID>
```
## Quick start for agents
零安装、零配置,复制一行就能用(只需要 [uv](https://docs.astral.sh/uv/)):
```bash
# 1) 直接读一篇公众号文章(不安装)
uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>
uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url> -f json
# 2) Claude Code:一行注册 MCP server(工具:fetch_wechat_article、parse_wechat_html;需要 Python 3.10+)
claude mcp add wechat-article-fetcher -- uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch-mcp
```
**Cursor**:写入 `~/.cursor/mcp.json`(全局)或项目里的 `.cursor/mcp.json`:
```json
{
"mcpServers": {
"wechat-article-fetcher": {
"command": "uvx",
"args": ["--from", "git+https://github.com/mrlong0129/wechat-article-fetcher", "wxfetch-mcp"]
}
}
}
```
**Agent Skill**(不装包也能用,Agent 会照着手册自己读):
```bash
# Claude Code
mkdir -p ~/.claude/skills/wechat-article-fetcher && curl -fsSL https://raw.githubusercontent.com/mrlong0129/wechat-article-fetcher/main/skill/SKILL.md -o ~/.claude/skills/wechat-article-fetcher/SKILL.md
# Cursor
mkdir -p ~/.cursor/skills/wechat-article-fetcher && curl -fsSL https://raw.githubusercontent.com/mrlong0129/wechat-article-fetcher/main/skill/SKILL.md -o ~/.cursor/skills/wechat-article-fetcher/SKILL.md
```
给编码 Agent 的完整说明见 [AGENTS.md](AGENTS.md)。要读微信以外的网页,用姊妹项目 [agent-web-fetch](https://github.com/mrlong0129/agent-web-fetch)。
## 功能
- 支持短链接 `https://mp.weixin.qq.com/s/<id>`、长链接 `...?__biz=...&mid=...&idx=...&sn=...`、不带 `https://` 的链接、从网页源码里复制出来带 `&` 的链接
- 一次多篇,或者从文件 / 标准输入批量读(会自动从一段文字里挑出公众号链接)
- 输出 Markdown(默认)或 JSON:标题、公众号名、作者、发布时间(北京时间 ISO 8601)、封面、摘要、正文 Markdown、纯文本、图片列表、用了哪种提取方式
- 被微信的「环境异常 / 去验证」页面挡住时自动换 User-Agent 重试;文章被删、违规、链接错误时直接说明原因
- 可选 `--download-images`:把图片存到本地并改写 Markdown 里的图片链接
- 依赖很少:`requests` + `beautifulsoup4` + `lxml`
## 安装
需要 Python 3.8+。
```bash
git clone https://github.com/mrlong0129/wechat-article-fetcher.git
cd wechat-article-fetcher
python3 -m venv .venv && source .venv/bin/activate
pip install -e . # 开发 / 跑测试:pip install -e '.[dev]'
wxfetch --help
```
## 用法
```bash
# 一篇文章 → Markdown 输出到终端
wxfetch https://mp.weixin.qq.com/s/<文章ID>
# JSON(单篇输出一个对象,多篇输出数组)
wxfetch https://mp.weixin.qq.com/s/<文章ID> -f json
# 多篇写进一个文件
wxfetch URL1 URL2 -o articles.md
# 从文件批量读(每行一个链接,# 开头的行忽略;也可以直接丢一段聊天记录进去)
wxfetch -i urls.txt --out-dir out/
# 每篇一个文件,并把图片下载到本地(out/<日期-标题>_images/)
wxfetch -i urls.txt --out-dir out/ --download-images
# Markdown 顶部加 YAML front matter(方便 Obsidian / 静态博客)
wxfetch URL --front-matter -o note.md
# 解析浏览器里「另存为」的 HTML,不联网
wxfetch --from-html saved_page.html -f json
# 查看抓取过程:用了哪个 UA、有没有被要求验证
wxfetch -v URL
```
| 参数 | 说明 |
| --- | --- |
| `-f markdown\|json` | 输出格式,默认 markdown |
| `-o FILE` / `--out-dir DIR` | 写进一个文件 / 每篇一个文件 |
| `-i FILE` | 从文件读链接,`-` 表示标准输入 |
| `--from-html FILE` | 解析本地 HTML,可重复 |
| `--download-images` / `--images-dir DIR` | 下载图片并改写链接 |
| `--front-matter` | Markdown 加 YAML 头 |
| `--ua ios\|android\|desktop` | 只用指定 UA(可重复,按顺序尝试) |
| `--delay 2.0` | 多篇之间的间隔秒数(带随机抖动) |
| `--retries 2` / `--timeout 20` | 每个 UA 的网络重试次数 / 单次超时 |
退出码:`0` 全部成功,`1` 至少一篇失败(错误信息在 stderr),`2` 用法错误。
### JSON 字段
下面是用测试里的示例页面(占位内容)跑出来的结果:
```jsonc
{
"url": "https://mp.weixin.qq.com/s/ExampleTextPost01",
"title": "这是一条用于测试的示例短内容,正文全部是占位文字。",
"account_nickname": "示例公众号", // 公众号名
"author": "示例作者", // 文章作者(可能和公众号名不同)
"publish_time": "2026-10-05T23:59:51+08:00",
"publish_timestamp": 1791215991,
"cover_image": "",
"description": "",
"content_markdown": "这是一条用于测试的示例短内容……",
"content_text": "这是一条用于测试的示例短内容……",
"images": [{"url": "https://mmbiz.qpic.cn/...", "alt": "", "local_path": "note_images/001.png"}],
"extraction_method": "js_content_noencode", // 哪种提取方式胜出
"extraction_candidates": {"js_content_noencode": 129, "og_meta": 129}, // 各方式提取到的字数
"canonical_url": "https://mp.weixin.qq.com/s/ExampleTextPost01",
"account_id": "gh_000000000000",
"account_alias": "ExampleAccount",
"item_show_type": 10, // 0 普通图文,10 纯文字短内容,8 图片消息
"fetched_with": "ios", // 哪个 UA 拿到的页面
"char_count": 140
}
```
### 作为 Python 库
```python
from wxfetch import fetch_article, extract_article
art = fetch_article("https://mp.weixin.qq.com/s/<文章ID>")
print(art.title, art.account_nickname, art.publish_time)
print(art.content_markdown)
art = extract_article(open("saved.html", encoding="utf-8").read()) # 离线解析
```
## How it works:怎么读一篇公众号文章
> 这一节写成了可以单独阅读的形式,不需要先看上面的用法说明。
### 问题:同一个链接,不同的人看到不同的页面
把一篇公众号文章的链接丢给一个通用的网页抓取服务,经常拿不到正文,只拿到一个写着「环境异常,完成验证后即可继续访问」的页面。可同一个链接在手机微信里点开一切正常。
原因很简单:网站会根据「你是谁、从哪来」决定给你什么。公众号文章本来就是给手机微信的内置浏览器看的;一个从数据中心 IP 发出、自称是某个爬虫的请求,在微信看来就很可疑,于是先丢一个验证页给你。
### 第一步:像手机微信那样去请求
每个 HTTP 请求都带一行 `User-Agent`,告诉网站「我是什么浏览器」。手机微信的 User-Agent 里有一段 `MicroMessenger/8.0.x`,网站就是靠它认出「这是微信」。
所以第一招就是把请求打扮成微信内置浏览器:iPhone 微信的 User-Agent,加上 `Referer: https://mp.weixin.qq.com/` 和 `Accept-Language: zh-CN`。实测公开文章的短链接用这一招基本都能直接拿到完整页面。拿不到就换 Android 微信,再不行换普通桌面 Chrome。
这不是破解。拿到的是任何人用微信打开这个公开链接都能看到的同一个页面,没有绕过登录、付费或者验证码。
### 第二步:先判断拿到的是什么页面
请求回来的不一定是文章,所以先分类:
- **验证页**:被跳转到 `/mp/wappoc_appmsgcaptcha`,或页面里有 `secitptpage`、「环境异常」「去验证」。这时换下一个 User-Agent,**不去尝试破解验证码**。
- **错误页**:页面里有一个 `weui-msg__title`,写着「该内容已被发布者删除」「此内容因违规无法查看」「参数错误」之类。这种情况换什么方式都没用,直接停下来,把原因告诉用户。
- **网络错误、限流(429)、服务端错误(5xx)**:同一种方式稍等再试,等待时间指数增长并带随机抖动。
批量读的时候每篇之间默认停 2 秒左右,别给对方添麻烦。
### 第三步:正文其实藏在好几个地方
拿到的页面有两三 MB,大部分是 JavaScript。正文通常不止一份:
1. **`<div id="js_content">`**:普通长文的正文就在这里。要用真正的 HTML 解析器去解析,不要用正则「从 js_content 截到 `<script>`」,否则会把一堆脚本当成正文。图片是懒加载的,真实地址在 `data-src`,`src` 里往往只是一个占位图。
2. **页面脚本里的 `content_noencode`**:纯文字的短内容(微信内部叫 `item_show_type = 10`)没有 `js_content`,正文以 JS 字符串的形式放在 `window.cgiDataNew` 里,换行写成 `\x0a`,尖括号写成 `\x3c`。把转义还原,就能得到带链接的正文。
3. **图片列表 `picture_page_info_list`**:图片消息(`item_show_type = 8`)的图片在这里。
4. **`og:title` 等分享元数据**:这是给分享卡片和搜索引擎预览用的 meta 标签。短内容没有标题,微信会把**整段正文**塞进 `og:title`,换行写成字面上的 `\n`。前三种都失败时,它是很好的兜底。
工具会把四种都试一遍,挑提取到字数最多的那份;字数差不多时优先结构最丰富的(`js_content` 有标题、列表、图片,`og:title` 只有纯文本),并记下是哪一种胜出,方便排查。
标题、公众号名、作者、发布时间这些元数据也一样有好几个来源(`var msg_title`、`var nickname`、`var ct` 这类 JS 变量,`cgiDataNew` 对象,页面元素,`og:` 标签),按顺序取第一个非空值。发布时间是 unix 时间戳,转成北京时间 `+08:00`。
### 第四步:把公众号排版变成干净的 Markdown
公众号的正文是层层嵌套的 `<section>` 和 `<span leaf>`,到处是内联样式。转换时只保留阅读需要的东西:标题、段落、列表、引用、加粗、链接、图片、代码块、简单表格。隐藏元素、公众号名片、小程序卡片、空段落都扔掉。
有两个和中文相关的坑:
- **加粗遇到标点**:按 CommonMark 规则,`他说**“重点”**然后` 渲染不出加粗,因为 `**` 紧挨着中文引号。解决办法是把两端的标点挪到外面:`他说“**重点**”然后`。
- **段内换行**:Markdown 里单个换行会被合并成一个空格。公众号里常见「一句一行」的写法,所以段内换行输出成 Markdown 的硬换行(行尾两个空格),段落之间仍然是空行。
### 图片的防盗链
公众号图片放在 `mmbiz.qpic.cn`。实测:不带 Referer,或者 Referer 是 `https://mp.weixin.qq.com/` 时返回原图;Referer 是别的网站时,会返回一张写着「此图片来自微信公众平台,未经允许不可引用」的 2KB 占位图。所以下载图片时带上公众号的 Referer。
### 什么时候会失效
- 长链接(`?__biz=...`)在微信客户端以外几乎都会被要求验证,短链接 `/s/<id>` 正常。遇到长链接,在微信里用「复制链接」拿短链接。
- 微信随时可能收紧检测,现在有效的 User-Agent 以后可能就不灵了。
- 这时更重的办法是用真实浏览器(比如 Playwright)打开页面。思路始终是:先用最便宜的方法,不行再上更重的。
## 局限
- **只能读公开文章。** 需要登录、付费、仅粉丝可见、已删除、违规屏蔽的文章都读不到,这个工具也不打算绕过这些限制。
- **长链接(`?__biz=...`)基本都会被要求验证。** 实测(2026-10)同一篇文章,短链接三种 UA 都能直接读,长链接不管用哪种 UA、哪个入口都会跳到验证页。
- **微信随时可能收紧检测。** 频繁请求、数据中心 IP 也更容易被要求验证。
- **不执行 JavaScript。** 完全靠脚本渲染的内容(视频号卡片、投票、部分互动组件)不会出现在结果里;视频只保留占位链接。
- 公众号排版五花八门,转出来的 Markdown 保证可读,不保证和原文排版一模一样。
- 计划中的兜底方案:可选的真实浏览器模式(如 [Playwright](https://playwright.dev/python/),作为可选依赖),纯 HTTP 被挡住时用无头浏览器打开页面。目前还没有实现,可以先在浏览器里打开文章、另存为 HTML,再用 `--from-html` 解析。
## 负责任地使用
- 文章版权归原作者和公众号所有。这个工具用于**个人阅读、笔记、学习研究**,请不要拿来批量搬运、转载或商用;转载请先获得授权并注明出处。
- 遵守微信公众平台的服务条款和当地法律法规。
- 控制频率:默认每篇间隔约 2 秒,失败会退避重试。请不要调得更激进,也不要大规模并发抓取。
- 被要求验证,说明对方不希望你这样访问。正确的做法是放慢速度、换个时间再试,或者回到微信 / 浏览器里正常阅读,而不是想办法破解验证码。
## 开发
```bash
pip install -e '.[dev]'
pytest # 离线测试(使用 tests/fixtures 里的页面)
pytest -m live # 可选:真实请求 mp.weixin.qq.com
```
`tests/fixtures/` 里都是精简过或合成的页面,不包含任何第三方文章正文:
- `text_post_trimmed.html`:按真实「纯文字短内容」页面(`item_show_type=10`)的结构精简而成,保留 `cgiDataNew`、`content_noencode`、`og:` 标签、`window.ct` 等关键结构和几个干扰项;正文、公众号名、ID 都换成了占位内容
- `verify_real.html`:真实的「环境异常」验证页(会话 token 已清除)
- `error_invalid_param_real.html`:真实的「参数错误」页
- `long_article_synthetic.html`、`verify_synthetic.html`、`deleted_synthetic.html`、`violation_synthetic.html`、`og_only_synthetic.html`:合成页面
更新记录见 [CHANGELOG.md](CHANGELOG.md)。
## English
**wxfetch** reads *public* WeChat Official Account articles (`mp.weixin.qq.com`) and outputs clean Markdown or JSON: title, account name, author, publish time (ISO 8601, `+08:00`), cover, description, Markdown body, plain text, image list, and which extraction method won.
```bash
pip install -e .
wxfetch https://mp.weixin.qq.com/s/<article-id> # Markdown to stdout
wxfetch URL1 URL2 -f json -o articles.json
wxfetch -i urls.txt --out-dir out/ --download-images
wxfetch --from-html saved_page.html # offline
```
**How it works.** Requests go out as WeChat's in-app browser (iPhone UA, then Android, then desktop Chrome) with a WeChat `Referer` and `Accept-Language: zh-CN`. Each response is classified first. The "环境异常 / 去验证" captcha page (`wappoc_appmsgcaptcha`, `secitptpage`) moves on to the next UA; nothing tries to solve the captcha. Deleted, removed or invalid-link pages stop with a clear reason. Network errors, 429 and 5xx are retried with exponential backoff. Content is extracted from up to four sources, and the richest one wins:
- the `#js_content` DOM, parsed with a real HTML parser, with lazy `data-src` images
- the `content_noencode` JS string, for text-only posts
- the picture list, for image posts
- the `og:title` / `og:description` meta tags; for short text posts WeChat stores the entire body in `og:title`, with literal `\n` escapes
The renderer keeps headings, lists, emphasis, links, images, code and tables. It moves punctuation outside `**` so that bold renders in CJK text, and writes line breaks inside a paragraph as Markdown hard breaks. Images are downloaded with `Referer: https://mp.weixin.qq.com/`, because a foreign Referer gets a placeholder image.
**Limitations.**
- Public articles only: no login, paid, fans-only or deleted content, and no attempt to bypass any of those.
- Outside WeChat, long `?__biz=` links are almost always challenged; use the short `/s/<id>` share link instead.
- WeChat may tighten detection at any time.
- No JavaScript execution. An optional headless-browser fallback (e.g. Playwright) is a possible future extra.
**Responsible use.** Articles are copyrighted by their authors. Use this tool for personal reading, notes and research. Respect WeChat's Terms of Service, keep request rates low (the default is about 2 s between articles), and don't republish content without permission.
## Related
- **[agent-web-fetch](https://github.com/mrlong0129/agent-web-fetch)**: a general web-reading toolkit for AI agents. It escalates only when blocked, reports why a page couldn't be read, and ships a CLI, an MCP server and an agent skill.
## Author & License
Made by Yomin Ma: [yomin.love](https://yomin.love) · [github.com/mrlong0129](https://github.com/mrlong0129)
[MIT](LICENSE) © 2026 Yomin Ma
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues