AgentLadle MCP CNINFO
# AgentLadle MCP CNINFO
**English** | [ไธญๆ](README_zh.md)
> ๐จ๐ณ/๐ญ๐ฐ Cloud-hosted MCP for A-share & HK listed companies (Past 3 years annual & latest interim reports). [Read more](https://github.com/achuan101/agentladle-mcp) | [Get API Key](https://agentladle.com/register)
A [MCP (Model Context Protocol)](https://modelcontextprotocol.io/) server that provides tools for **discovering, downloading, parsing, and searching** China A-share announcements from [CNINFO (ๅทจๆฝฎ่ต่ฎฏ็ฝ)](http://www.cninfo.com.cn).
It enables AI assistants (Claude, Cursor, etc.) to access CNINFO announcement data through 6 structured tools โ from discovering available announcements to keyword-searching within their pages.
> **Scope (v0.1):** Announcements only. Periodic reports (ๅนดๆฅ / ๅๅนดๆฅ / ไธๅญฃๆฅ / ไธๅญฃๆฅ) are out of scope.
## Features
- **6 MCP tools** for CNINFO announcement data: state-driven retrieval (search directly, fallback to download/parse only when needed)
- **PDF document parsing** using [PyMuPDF](https://pymupdf.readthedocs.io/) โ physical page extraction into page-split JSON
- **Local keyword search** with TF + position-boost scoring, zero external search dependencies
- **Idempotent** โ already-downloaded/parsed files are automatically skipped
- **Zero-config install** โ one line to add to your MCP client, no clone or manual setup needed
- **Pure Python**, cross-platform (Windows / macOS / Linux)
## Prerequisites
- **Python 3.10+** โ [Download Python](https://www.python.org/downloads/)
- **uv** โ [Install uv](https://docs.astral.sh/uv/getting-started/installation/)
> **Note:** After installing uv, restart your terminal and MCP client (e.g. Cherry Studio) to ensure the `uv` command is recognized.
## Quick Start
Add to your MCP client configuration (Claude Desktop, Cursor, etc.):
```json
{
"mcpServers": {
"mcp-cninfo": {
"command": "uvx",
"args": ["agentladle-mcp-cninfo"]
}
}
}
```
That's it. `uvx` will automatically download the package and its dependencies from PyPI โ no clone, no manual install, no path configuration.
### Alternative: pip install
If you prefer managing the environment yourself:
```bash
pip install agentladle-mcp-cninfo
```
Then configure:
```json
{
"mcpServers": {
"mcp-cninfo": {
"command": "agentladle-mcp-cninfo"
}
}
}
```
### Alternative: Run from source (local development)
Clone the repository and run directly:
```bash
git clone https://github.com/agentladle/mcp-cninfo.git
```
Then configure your MCP client:
```json
{
"mcpServers": {
"mcp-cninfo": {
"command": "uv",
"args": ["run", "--directory", "/path/to/mcp-cninfo", "agentladle-mcp-cninfo"]
}
}
}
```
Replace `/path/to/mcp-cninfo` with the actual path to the cloned repository.
## Data Flow
```
CNINFO API Local Files (~/.agentladle/mcp-cninfo/data/)
โโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
szse_stock.json โโโ companies.json (stock_codeโorgId mapping)
โ
hisAnnouncement/query โโโ pdf/{LOCAL_KEY}/ (Tool 2: primary PDF/HTML + manifest)
โ
PyMuPDF parsing โโโ json/*.json (Tool 3: parse, page-split)
โ
Local TF search โโโ search results (Tool 4: keyword search)
Page range read โโโ page content (Tool 5: read pages)
```
## Tools
| # | Tool | Description |
|---|------|-------------|
| 1 | `list_cninfo_announcements` | Discover available CNINFO announcements for a company |
| 2 | `download_cninfo_announcement` | Download announcement PDF (HTML fallback); idempotent |
| 3 | `parse_cninfo_announcement` | Parse PDF/HTML into page-split JSON using PyMuPDF |
| 4 | `keyword_search` | Full-text keyword search with TF relevance scoring |
| 5 | `get_announcement_pages` | Read announcement content by page number range |
| 6 | `lookup_stock_code` | **Diagnostic**: look up stock_codeโorgId mapping when resolution fails |
### Tool 1: `list_cninfo_announcements`
List available CNINFO announcements for a company. Use this tool ONLY when the exact date/title is unspecified by the user, or when a download attempt fails due to an ambiguous match. Default categories exclude periodic reports (ๅนดๆฅ / ๅๅนดๆฅ / ไธๅญฃๆฅ / ไธๅญฃๆฅ).
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `stock_code` | string | โ
| 6-digit stock code, e.g. `"000001"` |
| `category` | string | โ | Category key, short code, or Chinese label, e.g. `"่ฃไบไผ"`, `"DSH"`, `"category_dshgg_szsh"`. Omit to list default announcement categories |
| `start_date` | string | โ | Start date `YYYY-MM-DD` |
| `end_date` | string | โ | End date `YYYY-MM-DD` |
| `title_keyword` | string | โ | Title keyword filter |
| `limit` | int | โ | Max announcements to return, default 10, max 50 |
### Tool 2: `download_cninfo_announcement`
Download a specific CNINFO announcement from static.cninfo.com.cn. Prefer `local_key` from `list_cninfo_announcements` when available. Idempotent.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `stock_code` | string | โ
| 6-digit stock code, e.g. `"000001"` |
| `announce_date` | string | โ | Announce date `YYYY-MM-DD` (optional if `local_key` provided) |
| `title_keyword` | string | โ | Title substring to disambiguate same-day announcements |
| `category` | string | โ | Optional category filter |
| `announcement_id` | string | โ | CNINFO announcement id if known |
| `local_key` | string | โ | Exact local bundle key from list results |
### Tool 3: `parse_cninfo_announcement`
Parse a downloaded announcement PDF/HTML into page-split JSON. Uses PyMuPDF for PDF physical-page text extraction.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `local_key` | string | โ
| Bundle key returned by list/download, e.g. `"000001_DSH_2026-07-02_8b1ad607"` |
### Tool 4: `keyword_search`
Full-text keyword search across all pages. Results ranked by TF + position-boost score.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `local_key` | string | โ
| Bundle key |
| `keywords` | string[] | โ
| 1โ5 search keywords |
| `match_mode` | string | โ | `"ANY"` (default, any keyword matches) / `"ALL"` (all must match) |
| `max_results` | int | โ | Max results to return, default 5, max 50 |
### Tool 5: `get_announcement_pages`
Read full page content by page number range.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `local_key` | string | โ
| Bundle key |
| `start_page` | int | โ
| Start page number (1-based) |
| `page_count` | int | โ | Number of pages to return, default 3, max 5 |
### Tool 6: `lookup_stock_code`
Diagnostic tool: look up stock_codeโorgId mapping. Use only when `download_cninfo_announcement` / `list_cninfo_announcements` returns `Stock code not found`. Bypasses the session failed-code cache.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `stock_code` | string | โ
| 6-digit stock code, e.g. `"000001"` |
| `refresh` | bool | โ | Force re-download of `szse_stock.json` from CNINFO (default: `false`) |
## Configuration
On first run, a default config file is created at `~/.agentladle/mcp-cninfo/config.yaml`:
```yaml
paths:
data_dir: "~/.agentladle/mcp-cninfo/data"
pdf_dir: "~/.agentladle/mcp-cninfo/data/pdf"
json_dir: "~/.agentladle/mcp-cninfo/data/json"
download:
delay_between_requests: 0.3
min_file_size: 500
list_page_size: 30
list_max_pages: 5
company:
cache_ttl_days: 7
```
## Data Directory Structure
```
~/.agentladle/mcp-cninfo/
โโโ config.yaml # Configuration (auto-created)
โโโ data/
โโโ companies.json # stock_codeโorgId mapping (auto-downloaded & cached)
โโโ pdf/ # Downloaded announcement bundles
โ โโโ 000001_DSH_2026-07-02_8b1ad607/
โ โ โโโ primary.pdf
โ โ โโโ manifest.json
โ โโโ ...
โโโ json/ # Parsed page-split JSON
โโโ 000001_DSH_2026-07-02_8b1ad607.json
โโโ ...
```
**File naming convention:** `{STOCK_CODE}_{CAT_SHORT}_{ANNOUNCE_DATE}_{ID_HASH}`
## Example Usage
The tools are designed with an **EAFP (Easier to Ask for Forgiveness than Permission)** approach. AI assistants should attempt to retrieve data directly and rely on errors to trigger downloads.
**Scenario A: File already exists locally (Shortest Path)**
```
User: "Search 000001 board resolution for ๅ่ดญ"
1. keyword_search(local_key="000001_DSH_2026-07-02_8b1ad607", keywords=["ๅ่ดญ", "ๅณ่ฎฎ"])
โ Returns page snippets matching the keywords immediately.
```
**Scenario B: File missing (Fallback triggered)**
```
User: "What did Ping An Bank announce in its latest board notice?"
1. list_cninfo_announcements(stock_code="000001", category="่ฃไบไผ", limit=3)
โ Returns local_key / announce_date / title.
2. keyword_search(local_key="...", keywords=["่ฃไบไผ", "ๅณ่ฎฎ"])
โ Error: File not found.
3. download_cninfo_announcement(stock_code="000001", local_key="...")
โ Downloads PDF to ~/.agentladle/mcp-cninfo/data/pdf/
4. parse_cninfo_announcement(local_key="...")
โ Parses into JSON.
5. keyword_search(local_key="...", keywords=["่ฃไบไผ", "ๅณ่ฎฎ"])
โ Retries search and returns data.
```
## Tech Stack
| Component | Choice | Purpose |
|-----------|--------|---------|
| MCP Framework | `mcp` (FastMCP) | MCP server with stdio transport |
| HTTP Client | `httpx` | CNINFO API requests & file downloads |
| PDF Parsing | `pymupdf` + `beautifulsoup4` | PDF page text extraction; HTML fallback |
| Search | Python built-in | TF + position-boost scoring |
| Config | `pyyaml` | YAML configuration file |
## Project Structure
```
src/mcp_cninfo/
โโโ __init__.py
โโโ server.py # MCP Server entry point
โโโ config.py # Config loading (~/.agentladle/mcp-cninfo/config.yaml, singleton cached)
โโโ models.py # Data models
โโโ categories.py # Announcement category whitelist / blacklist
โโโ response.py # Unified JSON responses
โโโ instances.py # Service singletons
โโโ tools/
โ โโโ list_announcements.py # Tool 1: list_cninfo_announcements
โ โโโ download.py # Tool 2: download_cninfo_announcement
โ โโโ parse.py # Tool 3: parse_cninfo_announcement
โ โโโ search.py # Tool 4: keyword_search
โ โโโ page.py # Tool 5: get_announcement_pages
โ โโโ lookup.py # Tool 6: lookup_stock_code
โโโ services/
โโโ company.py # CNINFO szse_stock.json + stock_codeโorgId
โโโ downloader.py # CNINFO query + PDF download
โโโ parser.py # PDF/HTMLโJSON parsing (PyMuPDF)
โโโ searcher.py # Local JSON search + TF scoring
โโโ keys.py # local_key helpers
```
## License
MIT
TDQS
Scored across 6 tools
Each tool has a distinct role: listing, downloading, parsing, searching, page retrieval, and stock code lookup. No overlap in functionality, making it easy for an agent to select the correct tool.
All tool names follow a consistent verb_noun pattern in snake_case (e.g., list_cninfo_announcements, download_cninfo_announcement, lookup_stock_code). Minor variation in verb choice ('lookup' vs 'list') is negligible.
With 6 tools covering the core workflow of discovering, downloading, parsing, and searching announcements, the count is well-scoped and avoids bloat. Each tool serves a necessary step in the process.
The tool set covers the primary lifecycle: list, download, parse, keyword search within a document, and page retrieval. Missing a cross-announcement search or category listing, but these are minor gaps that do not break the workflow.