Stock-Tools MCP Server
README.md
# Stock-Tools MCP Server
An MCP (Model Context Protocol) server that gives any MCP-compatible AI agent — Claude Desktop, or any other MCP client — a full stock-analysis and ML toolkit: live market data, model training (scikit-learn + PyTorch), MLflow experiment tracking, a SQLite-backed model registry, champion/challenger promotion, and Docker containerization.
Built as a hands-on learning project to go from "what is MCP" to a working, containerized, MLOps-tracked agentic system — with all of the real debugging that implies.
## Key finding
The project trains two very different kinds of models, and the contrast between their results is the actual headline:
| Target | Result | Interpretation |
|---|---|---|
| Next-day **price direction** (up/down) | ~50–56% accuracy across multiple model types and feature sets | Consistent with the efficient market hypothesis — daily direction from price-only features carries essentially no learnable signal. This was confirmed independently three separate times using different approaches, all landing on the same conclusion. |
| Next-day **volatility regime** (high/low, tertile split) | **69–77% accuracy**, 7–25 points above each ticker's own majority-class baseline, across 5 tickers | Volatility clustering is a real, well-documented phenomenon (the basis of GARCH-family models used throughout quant finance) — and it shows up clearly here. |
**Volatility-regime results by ticker:**
| Ticker | Accuracy | Majority baseline | Lift |
|---|---|---|---|
| MSFT | 77.05% | 52.46% | +24.6 |
| AAPL | 69.35% | 51.61% | +17.7 |
| JPM | 67.74% | 54.84% | +12.9 |
| TSLA | 67.80% | 61.02% | +6.8 |
| ORCL | 60.32% | 53.97% | +6.3 |
The takeaway isn't "this predicts stock prices" — it's that the project correctly identified *which* target has real signal and which doesn't, and built the evaluation rigor (chronological splits, majority-class baselines, MLflow tracking) to prove both findings rather than just assert them.
## Architecture
```mermaid
flowchart LR
A[MCP Client<br/>e.g. Claude Desktop] -- MCP over stdio --> B[stock-tools MCP Server]
B --> C[yfinance<br/>market data]
B --> D[scikit-learn / PyTorch<br/>model training]
B --> E[MLflow<br/>experiment tracking]
B --> F[SQLite<br/>experiment log + production registry]
G[Docker container] -.wraps.-> B
```
The server runs as a standard MCP stdio server. It can run directly in a Python virtual environment, or containerized via Docker — both are supported and documented below.
## Tools exposed
| Tool | Description |
|---|---|
| `get_stock_history(ticker, period)` | Historical closing prices for a ticker |
| `train_model(ticker)` | Trains a logistic regression next-day **direction** classifier |
| `train_lstm_model(ticker)` | Trains a PyTorch LSTM next-day **direction** classifier |
| `predict_next_move(ticker)` | Predicts tomorrow's direction using the latest trained model |
| `train_volatility_model(ticker)` | Trains a logistic regression **volatility-regime** classifier (tertile split, true range + volume features) |
| `retrain_and_promote(ticker)` | Champion/challenger pattern — trains a fresh volatility model and promotes it to "production" only if it beats the current one |
| `list_trained_models()` | Lists every model trained so far, with accuracy |
| `list_production_models()` | Lists current production models per ticker |
Every training run is logged to MLflow (`stock-direction` and `stock-volatility` experiments) and to a SQLite metadata table, so results are always comparable and auditable rather than one-off numbers.
## Tech stack
All free/open-source, runs entirely locally — no paid services, no cloud accounts required.
| Purpose | Tool |
|---|---|
| Market data | `yfinance` |
| Classical ML | scikit-learn |
| Deep learning | PyTorch |
| Experiment tracking | MLflow (SQLite backend) |
| Metadata / registry | SQLite |
| MCP server framework | `mcp` Python SDK (`MCPServer`) |
| Containerization | Docker |
| Local testing (no LLM needed) | MCP Inspector |
## Setup
### Prerequisites
- Python 3.11+
- [Docker Desktop](https://www.docker.com/products/docker-desktop/) (for the containerized workflow)
- [Claude Desktop](https://claude.ai/download) (or any other MCP-compatible client)
- Node.js (only needed for `npx`, to run the MCP Inspector)
### 1. Clone and set up the environment
```powershell
git clone https://github.com/Osmium-hacked/stock-tools-mcp.git
cd stock-tools-mcp
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt
```
### 2. Test locally with the MCP Inspector (no LLM required)
This is the fastest way to verify the server works before wiring it into any AI client:
```powershell
npx @modelcontextprotocol/inspector "venv\Scripts\python.exe" server.py
```
Opens a browser UI where you can call each tool directly and inspect the raw response.
### 3. Connect to Claude Desktop
Open Claude Desktop's config file (Settings → Developer → Edit Config) and add:
```json
{
"mcpServers": {
"stock-tools": {
"command": "C:\\path\\to\\stock-tools-mcp\\venv\\Scripts\\python.exe",
"args": ["C:\\path\\to\\stock-tools-mcp\\server.py"]
}
}
}
```
Fully quit Claude Desktop (system tray → Exit, not just closing the window) and reopen it. The tools should appear under `stock-tools` in a new chat.
### 4. (Optional) Run containerized instead
```powershell
docker build -t stock-tools:latest .
```
Update the Claude Desktop config to launch the container instead of the venv directly:
```json
{
"mcpServers": {
"stock-tools": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-v", "C:\\path\\to\\stock-tools-mcp\\data:/app/data",
"stock-tools:latest"
]
}
}
}
```
Note: `-i`, not `-it` — Claude Desktop pipes raw JSON-RPC over stdio, and the `-t` pseudo-terminal flag would corrupt that stream. Also: any time you edit `server.py`, you need to `docker build` again before restarting Claude Desktop — the image is a frozen snapshot taken at build time, not a live view of your files.
### 5. View experiment tracking
```powershell
mlflow ui --backend-store-uri sqlite:///C:/path/to/stock-tools-mcp/data/mlflow.db
```
Opens a dashboard at `http://localhost:5000` with every training run, comparable side by side.
## Usage examples
Once connected, just talk to your MCP client naturally:
- *"What's AAPL done over the past month?"*
- *"Use train_volatility_model to train a volatility model for MSFT"*
- *"Use retrain_and_promote for TSLA"*
- *"Use list_production_models to show everything currently in production"*
## Project structure
```
stock-tools-mcp/
├── server.py # MCP server + all tool definitions
├── requirements.txt # Hand-curated (not a raw pip freeze — see Lessons Learned)
├── Dockerfile
├── .gitignore
├── README.md
└── data/ # gitignored — created at runtime
├── models/ # trained model artifacts (.joblib, .pt)
├── experiments.db # training run metadata
└── mlflow.db # MLflow tracking backend
```
## Lessons learned
A few things worth being upfront about, since they shaped the project as much as the results did:
- **Daily price direction is close to a random walk.** This was confirmed three separate ways (different model architectures, different feature sets) before accepting it rather than continuing to chase a better number. Efficient markets are a real constraint, not a modeling failure to be tuned away.
- **The first volatility-model attempt (52% accuracy) looked like another dead end — it wasn't.** A median-split label meant most days landed in the ambiguous middle of the distribution and got an arbitrary label. Switching to a tertile split (dropping the ambiguous middle third entirely) revealed the real signal that was there the whole time. Worth remembering: a "negative result" is sometimes a label design problem, not a ceiling.
- **`requirements.txt` should be hand-curated, not a raw `pip freeze`.** A frozen Windows venv includes OS-specific transitive dependencies (e.g. `pywin32`) that don't exist on Linux and will break a Docker build. Listing only what the code directly imports lets pip resolve correct platform-specific versions itself.
- **Relative paths break under a subprocess launcher.** Claude Desktop (and Docker) don't guarantee the working directory you assume. Every path in `server.py` is anchored to `Path(__file__).resolve().parent`, not the current working directory, after this caused two separate crashes.
- **Small-sample caveat:** the tertile split shrinks the effective test set further. The volatility results are a real, consistent, multi-ticker finding — but with 5 tickers and one static train/test split each, they're a strong signal to build on, not a number to over-claim in production.
## Future work
- Walk-forward / rolling-origin backtesting instead of a single static split
- Richer volatility features (implied vol if a free options-data source is found, cross-asset signals like VIX)
- Expand `retrain_and_promote` into a genuine two-agent system (a Trainer role and an independent Reviewer role that must separately approve promotion)
- Extend the production registry to serve predictions directly, not just track training runs
## License
MITThis server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues