MarketSage
Provides access to Hugging Face finance datasets and optional models, enabling dataset status checks and sentiment scoring with FinBERT or the deterministic fallback.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MarketSagePull a market snapshot and recent sentiment for AAPL"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MarketSage
An MCP-native market intelligence workbench: a traditional analyst workflow, exposed to LLM clients as tools, with every run saved and every source caveated.
A Go MCP gateway exposes seven finance tools and a saved-run resource over stdio. Behind it, a Python FastAPI analytics core provides market snapshots, price history, sentiment scoring, evidence search and research briefs, with an OpenBB-ready market adapter and a Hugging Face dataset boundary. DuckDB persists dataset manifests, research runs and audit events. A Next.js workbench gives analysts the same capabilities as a conventional application surface.
MarketSage does not execute trades, move money or provide investment advice.
At a glance
The problem | Analyst teams have a workflow that works. LLM clients want to use it. Bolting a chat box onto a finance app gives the model no structure, no provenance and no record; exposing the workflow as typed tools with saved, auditable runs does. |
What it does | Seven MCP tools ( |
Stack | Go 1.24 (MCP gateway), Python 3.12 with FastAPI and |
Validation | One gate, |
Related MCP server: repo-explorer-mcp
Results
Every figure below was observed by npm run eval, which runs offline with a fixed seed and no API key, and writes metrics/headline.json. Committed FinanceBench and FiQA fixtures replayed through the analytics core with the deterministic providers bound; no network, no key, no model download. Rows marked pending need hardware, data or a service the offline harness does not have; nothing here is estimated.
Metric | Value | How it was measured |
Evidence recall@5, ticker given | 94.7% | BM25 over 145 FinanceBench filing excerpts, 150 labelled queries; without the ticker hint recall@5 is 44.0% (MRR 0.33, nDCG@10 0.38), so the structured hint does the heavy lifting |
Sentiment fallback commits on | 8.1% | marketsage-lexicon-v0 matched 19 of 234 FiQA test sentences and was right on 89.5% of those; accuracy over every row 8.1% against a 61.5% majority baseline; ECE 0.24; FinBERT is the real path and is pending |
Brief claims grounded | 12 / 12 | claim bullets whose reference resolves to a snapshot, evidence id or sentiment hash in the same payload; 5 tickers, 3 with no committed evidence said so instead of borrowing another company's |
Responses conforming to contract | 8 / 8 | every tool response validated against the JSON Schema generated from the Pydantic models; Go decodes the same fixtures with unknown fields forbidden; schema file in sync |
Fault injections handled honestly | 9 / 9 | named scenarios that degraded with an explicit warning or failed with a clean error and an audit row; none returned invented data |
Tool calls with an audit row | 8 / 8 | replayed demo chain; 8 ok rows, 0 orphans, 8/8 envelopes carry source, mode, timestamp and caveats |
FinanceBench recall by scorer and hint (150 queries)
BM25, ticker given @1 |
| 47.3% |
BM25, ticker given @5 |
| 94.7% |
BM25, ticker given @10 |
| 100.0% |
BM25, no hint @1 |
| 20.7% |
BM25, no hint @5 |
| 44.0% |
BM25, no hint @10 |
| 60.7% |
Term overlap, no hint @1 |
| 12.0% |
Term overlap, no hint @5 |
| 26.7% |
Term overlap, no hint @10 |
| 38.0% |
Observed offline, and what is not
Status | Evidence | |
Committed fixtures | observed | FinanceBenchRetrieval corpus/test 145 rows, FinanceBenchRetrieval queries/test 150 rows, FinanceBenchRetrieval qrels/test 150 rows, fiqa-sentiment-classification default/test 234 rows; MIT licensed, sha256-pinned in data/fixtures/manifest.json |
Sentiment confidence calibration | observed | ECE 0.24 on 19 committed predictions; the confidence field is a term-count formula, not a probability |
Evidence coverage | observed | 2 of 5 seeded tickers have filing excerpts in the corpus; the rest get an explicit coverage warning |
Latency budget (NFR-008) | observed | 8 of 8 tools under 2 s p95 in-process over 5 runs; timings kept in metrics/eval-latest.json, not here, because they vary by host |
MCP surface | observed | 7 tools, 4 prompts, 1 resource template, read from |
FinBERT sentiment accuracy | pending | pending: needs MARKETSAGE_ENABLE_MODEL_DOWNLOADS=true and a model download; the offline harness scores the lexicon fallback only |
Embedding retrieval | pending | pending: bge-small / MiniLM path is not implemented; lexical scorers only |
Live OpenBB market data | pending | pending: seeded prices are illustrative and are not scored; live mode needs the optional OpenBB dependency and provider configuration |
Architecture
LLM host / MCP client ──stdio JSON-RPC──► Go MCP gateway ──HTTP──► Python analytics core
├── OpenBB-ready market adapter
Next.js analyst workbench ──server-side proxy───────────────────────────────────► ├── Hugging Face dataset/model boundary
└── DuckDB: manifests, runs, audit eventsThree languages, each where it is strongest: Go for a transport-disciplined MCP server that never writes to stdout in stdio mode, Python for the data and model integrations, TypeScript for the client and the product surface. One schema in packages/contracts/marketsage.schema.json describes the shared payloads. It is generated from the Pydantic models, every response is validated against it, and the Go client decodes the exported fixtures with unknown fields forbidden.
Quick start
Requires Node.js 22+, Go 1.24+ and uv.
npm install
npm run check # the gate
npm run demo:mcp # a TypeScript MCP client starts the stack, lists tools, runs the chain, reads a saved run
npm run eval # the offline evaluation behind the results card above; npm run card re-renders itFor the workbench, in two terminals:
npm run dev:analytics
npm run dev --workspace apps/web # http://localhost:3000, then Run BriefData and model modes
Mode | Behaviour |
| Deterministic local data from |
| Tries live OpenBB data and falls back to seeded data, with a warning in the response so the fallback is never silent. |
| Requires the optional OpenBB dependencies and fails clearly when they are missing. |
Model downloads are off by default. MARKETSAGE_ENABLE_MODEL_DOWNLOADS=true enables FinBERT; otherwise the deterministic sentiment fallback is used and reported as such.
Protected local mode
The analytics API is open for local demos. Set MARKETSAGE_HTTP_TOKEN to require bearer auth; the Go gateway and the Next.js proxy forward the same token server-side.
Documentation
The problem, the design and its reasons, what is measured | |
A guided tour of every feature, with commands and files | |
A five-minute walkthrough | |
Security posture, dependency sweeps, operational notes | |
Which datasets and models were reviewed, and why some were excluded | |
Licences of everything used | |
Validation evidence and known limitations | |
What the evaluation harness measures, how, and what it deliberately leaves pending | |
Requirements, high-level design, low-level design, execution plan, decisions | |
The codebase knowledge graph (graphify): how to build it, query it, and what it excludes |
License
AGPL-3.0-only, because OpenBB is. To relicense permissively, isolate OpenBB behind an external service boundary first and confirm compatibility.
This server cannot be deployed
Maintenance
Related MCP Connectors
Financial data MCP server for Claude, ChatGPT, Cursor and Codex. Real-time stock quotes, financial statements, options flow, SEC filings, insider trades, 13F holdings, macro data and market news from gloom.sh, the open-source Bloomberg Terminal alternative.
Unified financial infrastructure connecting AI agents directly to trade live/demo brokerage accounts, Web3 non-custodial wallets, real-time market data across equities, ETFs, crypto, forex, options, DeFi swaps, and prediction markets, institutional research feeds, and algorithmic strategy backtesters.
Save and query market signals from your AI conversations.
Evidence-backed crypto due diligence with sources, freshness, and a runtime receipt on every call.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceEnables EHR Copilot operations such as order cloning, queue building, and execution trace analysis over stdio.-
- FlicenseNot gradedqualityAmaintenanceEnables LLM-driven repository exploration by exposing an explore_repository tool over stdio, using codebase memory and ripgrep/rtk search to navigate and analyze codebases.-
- FlicenseNot gradedqualityCmaintenanceEnables MCP hosts like Cursor or Claude to query an enterprise knowledge base via stdio, providing RAG-based answers with cited sources.-
- AlicenseBqualityBmaintenanceEnables IDE and DeepSeek Harness agents to control an agent-kernel instance over stdio, exposing tools for auth, projects, assignments, runs, scheduler nudges, and executor settings.12MIT