agentladle/mcp-sec
# AgentLadle MCP SEC
**English** | [δΈζ](README_zh.md) | πΊ [Watch Demo](https://www.youtube.com/watch?v=qZteRG7WvIQ)
> π¨π³/ππ° Cloud-hosted MCP for A-share & HK listed companies (Past 3 years annual & latest interim reports). [Read more](https://github.com/achuan101/agentladle-mcp) | [Get API Key](https://agentladle.com/register)
A [MCP (Model Context Protocol)](https://modelcontextprotocol.io/) server that provides tools for **discovering, downloading, parsing, and searching** U.S. SEC financial reports.
It enables AI assistants (Claude, Cursor, etc.) to access SEC EDGAR data through 6 structured tools β from discovering available filings to keyword-searching within their pages.
## Features
- **6 MCP tools** for SEC financial data: state-driven retrieval (search directly, fallback to download/parse only when needed)
- **Professional SEC document parsing** using [edgartools](https://github.com/dgunning/edgar-tools) β accurate page-break detection and structured node-tree extraction for iXBRL filings
- **Local keyword search** with TF + position-boost scoring, zero external dependencies
- **Idempotent** β already-downloaded/parsed files are automatically skipped
- **Zero-config install** β one line to add to your MCP client, no clone or manual setup needed
- **Pure Python**, cross-platform (Windows / macOS / Linux)
## Prerequisites
- **Python 3.10+** β [Download Python](https://www.python.org/downloads/)
- **uv** β [Install uv](https://docs.astral.sh/uv/getting-started/installation/)
> **Note:** After installing uv, restart your terminal and MCP client (e.g. Cherry Studio) to ensure the `uv` command is recognized.
## Quick Start
Add to your MCP client configuration (Claude Desktop, Cursor, etc.):
```json
{
"mcpServers": {
"mcp-sec": {
"command": "uvx",
"args": ["agentladle-mcp-sec"],
"env": {
"SEC_EMAIL": "your@email.com"
}
}
}
}
```
That's it. `uvx` will automatically download the package and its dependencies from PyPI β no clone, no manual install, no path configuration.
> β οΈ **SEC Email Requirement:** Replace `your@email.com` with your real email. The SEC requires a valid email in the User-Agent header. Using a fake email may result in your IP being blocked.
### Alternative: pip install
If you prefer managing the environment yourself:
```bash
pip install agentladle-mcp-sec
```
Then configure:
```json
{
"mcpServers": {
"mcp-sec": {
"command": "agentladle-mcp-sec",
"env": {
"SEC_EMAIL": "your@email.com"
}
}
}
}
```
### Alternative: Run from source (local development)
Clone the repository and run directly:
```bash
git clone https://github.com/agentladle/mcp-sec.git
```
Then configure your MCP client:
```json
{
"mcpServers": {
"mcp-sec": {
"command": "uv",
"args": ["run", "--directory", "/path/to/mcp-sec", "agentladle-mcp-sec"],
"env": {
"SEC_EMAIL": "your@email.com"
}
}
}
}
```
Replace `/path/to/mcp-sec` with the actual path to the cloned repository.
## Data Flow
```
SEC EDGAR API Local Files (~/.agentladle/mcp-sec/data/)
ββββββββββββββ ββββββββββββββββββββββββββββββ
company_tickers.json βββ company_tickers.json (tickerβCIK mapping)
β
SEC Submissions API βββ html/{TICKER_FORM_DATE}/ (Tool 2: primary + HTML exhibits)
β
edgartools parsing βββ json/*.json (Tool 3: parse, page-split)
β
Local TF search βββ search results (Tool 4: keyword search)
Page range read βββ page content (Tool 5: read pages)
TOC lookup βββ table of contents (Tool 6: get TOC)
```
## Tools
| # | Tool | Description |
|---|------|-------------|
| 1 | `list_sec_filings` | Discover available SEC filings for a company |
| 2 | `download_sec_report` | Download SEC filing; 6-K/8-K include HTML exhibits by default (PDF skipped) |
| 3 | `parse_sec_report` | Parse HTML into page-split JSON using edgartools |
| 4 | `keyword_search` | Full-text keyword search with TF relevance scoring |
| 5 | `get_report_pages` | Read report content by page number range |
| 6 | `get_report_toc` | Get the Table of Contents page(s) |
| 7 | `lookup_ticker_cik` | **Diagnostic**: look up tickerβCIK mapping when CIK resolution fails |
### Tool 1: `list_sec_filings`
List available SEC filings for a company. Use this tool ONLY when the exact year/date is unspecified by the user, or when a download attempt fails due to an invalid date.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `ticker` | string | β
| Stock ticker, e.g. `"AAPL"` |
| `form` | string | β | Filing type filter, e.g. `"10-K"`. Omit to list all financial report types (10-K, 10-Q, 20-F, 6-K, 8-K, 40-F) |
| `limit` | int | β | Max filings to return, default 5, max 20 |
### Tool 2: `download_sec_report`
Download a specific SEC filing from EDGAR. For **6-K/8-K**, also downloads **HTML exhibits** by default (PDF exhibits are skipped, not parsed). Idempotent.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `ticker` | string | β
| Stock ticker, e.g. `"AAPL"` |
| `form` | string | β
| Filing type: `"10-K"`, `"10-Q"`, `"20-F"`, `"6-K"`, `"8-K"` |
| `report_date` | string | β
| Report date (fiscal period end date), e.g. `"2025-01-31"` |
| `include_exhibits` | bool | β | Download HTML exhibits. Default: `true` for 6-K/8-K, otherwise `false`. PDFs are never parsed |
### Tool 3: `parse_sec_report`
Parse a downloaded HTML filing into page-split JSON. Uses edgartools `mark_page_breaks()` + `parse_html()` for professional SEC document parsing.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `ticker` | string | β
| Stock ticker |
| `form` | string | β
| Filing type |
| `report_date` | string | β
| Report date (fiscal period end date) |
### Tool 4: `keyword_search`
Full-text keyword search across all pages. Results ranked by TF + position-boost score.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `ticker` | string | β
| Stock ticker |
| `form` | string | β
| Filing type |
| `report_date` | string | β
| Report date (fiscal period end date) |
| `keywords` | string[] | β
| 1β5 search keywords |
| `match_mode` | string | β | `"ANY"` (default, any keyword matches) / `"ALL"` (all must match) |
| `max_results` | int | β | Max results to return, default 5, max 50 |
### Tool 5: `get_report_pages`
Read full page content by page number range.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `ticker` | string | β
| Stock ticker |
| `form` | string | β
| Filing type |
| `report_date` | string | β
| Report date (fiscal period end date) |
| `start_page` | int | β
| Start page number (1-based) |
| `page_count` | int | β | Number of pages to return, default 3, max 5 |
### Tool 6: `get_report_toc`
Get the Table of Contents page(s). Searches the first 10 pages for "Table of Contents".
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `ticker` | string | β
| Stock ticker |
| `form` | string | β
| Filing type |
| `report_date` | string | β
| Report date (fiscal period end date) |
### Tool 7: `lookup_ticker_cik`
Diagnostic tool: look up tickerβCIK mapping. Use only when `download_sec_report` / `list_sec_filings` returns `CIK not found` or `Ticker not found`. Bypasses the session failed-ticker cache and returns same-CIK alias tickers.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `ticker` | string | β
| Stock ticker, e.g. `"BABA"` |
| `refresh` | bool | β | Force re-download of `company_tickers.json` from SEC (default: `false`) |
## Configuration
On first run, a default config file is created at `~/.agentladle/mcp-sec/config.yaml`:
```yaml
sec:
email: ""
paths:
data_dir: "~/.agentladle/mcp-sec/data"
html_dir: "~/.agentladle/mcp-sec/data/html"
json_dir: "~/.agentladle/mcp-sec/data/json"
download:
delay_between_requests: 0.2
min_file_size: 5000
```
The `email` field is used to build the SEC-compliant User-Agent header (`AgentLadleMcpSec {email}`). You can configure it in three ways (in order of priority):
1. **Environment variable** `SEC_EMAIL` β recommended, set it in your MCP client JSON config
2. **Config file** β edit `~/.agentladle/mcp-sec/config.yaml` and set `email`
3. **Default** β if empty, a placeholder email is used (not recommended for production)
> β οΈ **SEC User-Agent Policy**: The SEC requires a real email in the User-Agent header. Using the default placeholder may result in your IP being blocked and can cause intermittent tickerβCIK lookup failures. `SEC_EMAIL` is required β please configure it.
## Data Directory Structure
```
~/.agentladle/mcp-sec/
βββ config.yaml # Configuration (auto-created)
βββ data/
βββ company_tickers.json # tickerβCIK mapping (auto-downloaded & cached)
βββ html/ # Downloaded HTML filings
β βββ AAPL_10-K_2025-01-31.htm
β βββ ...
βββ json/ # Parsed page-split JSON
βββ AAPL_10-K_2025-01-31.json
βββ ...
```
**File naming convention:** `{TICKER}_{FORM}_{REPORT_DATE}.htm/json`
## Example Usage
The tools are designed with an **EAFP (Easier to Ask for Forgiveness than Permission)** approach. AI assistants should attempt to retrieve data directly and rely on errors to trigger downloads.
**Scenario A: File already exists locally (Shortest Path)**
```
User: "Analyze AAPL's latest 10-K management discussion"
1. keyword_search(ticker="AAPL", form="10-K", report_date="2025-01-31", keywords=["management", "discussion"])
β Returns page snippets matching the keywords immediately.
```
**Scenario B: File missing (Fallback triggered)**
```
User: "What is Tesla's 2024 revenue?"
1. keyword_search(ticker="TSLA", form="10-K", report_date="2024-12-31", keywords=["revenue", "net sales"])
β Error: File not found.
2. download_sec_report(ticker="TSLA", form="10-K", report_date="2024-12-31")
β Downloads HTML to ~/.agentladle/mcp-sec/data/html/
3. parse_sec_report(ticker="TSLA", form="10-K", report_date="2024-12-31")
β Parses into JSON.
4. keyword_search(ticker="TSLA", form="10-K", report_date="2024-12-31", keywords=["revenue", "net sales"])
β Retries search and returns data.
```
## Tech Stack
| Component | Choice | Purpose |
|-----------|--------|---------|
| MCP Framework | `mcp` (FastMCP) | MCP server with stdio transport |
| HTTP Client | `httpx` | SEC API requests & file downloads |
| HTML Parsing | `edgartools` + `beautifulsoup4` | Professional SEC iXBRL parsing (page-break detection + node tree) |
| Search | Python built-in | TF + position-boost scoring |
| Config | `pyyaml` | YAML configuration file |
## Project Structure
```
src/mcp_sec/
βββ __init__.py
βββ server.py # MCP Server entry point
βββ config.py # Config loading (~/.agentladle/mcp-sec/config.yaml, singleton cached)
βββ models.py # Data models
βββ tools/
β βββ list_filings.py # Tool 1: list_sec_filings
β βββ download.py # Tool 2: download_sec_report
β βββ parse.py # Tool 3: parse_sec_report
β βββ search.py # Tool 4: keyword_search
β βββ page.py # Tool 5: get_report_pages
β βββ toc.py # Tool 6: get_report_toc
βββ services/
βββ downloader.py # SEC EDGAR download + tickerβCIK
βββ parser.py # HTMLβJSON parsing (edgartools)
βββ searcher.py # Local JSON search + TF scoring
```
## License
MITTDQS
Scored across 7 tools
Each tool targets a distinct stage of the SEC report lifecycle: discovery (list_sec_filings, lookup_ticker_cik), acquisition (download_sec_report), processing (parse_sec_report), and content access (keyword_search, get_report_pages, get_report_toc). The detailed strategy notes explicitly separate search vs. page reading vs. TOC retrieval, so an agent can reliably select the right tool.
Six of seven tools follow a clear verb_noun pattern (download_, parse_, get_, list_, lookup_), but keyword_search breaks the convention by leading with a noun/modifier instead of a verb (e.g., search_keywords would be consistent). Overall the pattern is still predictable and readable.
Seven tools is well-scoped for the server's purpose: a complete pipeline from ticker resolution to full-text search and page retrieval. Each tool covers a distinct function without redundancy, and the count sits comfortably in the ideal 3-15 range.
The tooling covers the full lifecycle of SEC report analysis: listing available filings, resolving tickers, downloading, parsing, searching, navigating via TOC, and reading specific page ranges. There are no obvious dead ends; the explicit fallback flow from error-prone tools to download/parse/lookup ensures agents can recover.