actian-docs-mcp
# docs-mcp




An [MCP (Model Context Protocol)](https://modelcontextprotocol.io) server that exposes Actian product documentation, GA4 content analytics, and Jenkins CI build status to Claude, enabling natural-language queries across all three sources in a single conversation.
> **Runs with no credentials.** The corpus, GA4 analytics, and Jenkins data ship as realistic mock fixtures, so you can clone, build, and query it in about two minutes. Every source has a documented upgrade path to live data. See [Live data](#live-data).
```
Claude Desktop / Claude Code
│
│ JSON-RPC over stdio
▼
docs-mcp
│
├── search_docs → Actian Markdown corpus (TF-IDF, swap for pgvector)
├── get_topic → Full topic content by ID
├── list_topics → Corpus index with product/type filters
├── get_content_gaps → GA4 zero-result search queries (mock → BigQuery)
└── get_build_status → Jenkins publish pipeline status (mock → live API)
```
## What this enables
Once connected, you can ask Claude:
> *"What are the top 10 search queries on the Actian docs site that returned no results this month?"*
> *"Find every topic in the Analytics Engine 8.0 corpus that mentions the Tableau connector, and check whether the publish pipeline is currently green."*
> *"Did the docs build pass? If it failed, search the corpus for the RPM upgrade procedure and show me the relevant steps."*
## Demo
<!-- TODO: record a short GIF of Claude calling get_content_gaps and returning
a ranked list, save it as docs/demo.gif, then uncomment the line below.
This is the single highest-value addition to this README. -->
<!--  -->
---
## Quick start
### 1. Clone and install
```bash
git clone https://github.com/Bipin-24/docs-mcp.git
cd docs-mcp
npm install
npm run build
```
### 2. Connect to Claude Desktop
Open your Claude Desktop config file:
| OS | Path |
|----|------|
| macOS | `~/Library/Application Support/Claude/claude_desktop_config.json` |
| Windows | `%APPDATA%\Claude\claude_desktop_config.json` |
Add this block (replace the path with your actual clone location):
```json
{
"mcpServers": {
"actian-docs": {
"command": "node",
"args": ["/absolute/path/to/docs-mcp/dist/index.js"]
}
}
}
```
Restart Claude Desktop. You should see `actian-docs` in the tools list.
### 3. Connect to Claude Code
Drop a `.mcp.json` file in your project root:
```json
{
"mcpServers": {
"actian-docs": {
"command": "node",
"args": ["../docs-mcp/dist/index.js"]
}
}
}
```
---
## Tools
### `search_docs`
Keyword search across the documentation corpus, ranked by TF-IDF with title and tag match bonuses. Upgradeable to embedding-based semantic search. See [Architecture](#search).
```
Input:
query string required Natural language search query
product string optional analytics-engine | ingres | actian-client | all
version string optional e.g. "8.0", "11.x"
limit number optional 1–10, default 5
Output: ranked list of matching topics with excerpt and relevance score
```
**Example prompt:** *"How do I upgrade Analytics Engine using RPM packages?"*
---
### `get_topic`
Retrieve the full Markdown content of a topic by ID.
```
Input:
topic_id string required Topic ID from search_docs results
Output: full topic content with metadata
```
---
### `list_topics`
Browse the corpus index.
```
Input:
product string optional analytics-engine | ingres | actian-client | all
topic_type string optional concept | task | reference | troubleshooting | all
Output: topics grouped by product
```
---
### `get_content_gaps`
Surfaces search queries that returned zero or very few results, meaning documentation your users need but that does not exist yet.
```
Input:
days number optional Lookback window, default 30
limit number optional Max gaps to return, default 20
min_searches number optional Minimum search volume, default 2
Output: ranked gap list with gap type, search volume, and nearest existing topics
```
**Example prompt:** *"What are the top missing documentation topics based on search data from the last 30 days?"*
---
### `get_build_status`
Returns the current Jenkins CI pipeline status for the docs publish job.
```
Input:
job string optional Jenkins job name, default "actian-docs-publish"
Output: latest build result, stage breakdown, and recent history
```
---
## Architecture
### Search
The server ships with a lightweight TF-IDF keyword search engine (`src/lib/search.ts`) that requires no external dependencies or API keys. It scores documents using term frequency with title and tag match bonuses.
To upgrade to embedding-based semantic search:
1. Add `chromadb` to `package.json`
2. Run `npm run index` to embed the corpus using the Python indexer (`scripts/index_corpus.py`)
3. Swap `scoreTopics()` in `src/lib/search.ts` for a Chroma similarity query
### Live data
The server ships with realistic mock data for GA4 analytics and Jenkins. To switch to live sources, add credentials to `.env`:
```bash
cp .env.example .env
# Edit .env with your BigQuery project ID and Jenkins token
```
See `.env.example` for all available configuration options.
The BigQuery query for GA4 Site Search export is in `scripts/ga4_export.sql`.
---
## Project structure
```
docs-mcp/
├── src/
│ ├── index.ts # MCP server — tool registry and routing
│ ├── tools/
│ │ ├── searchDocs.ts # search_docs handler
│ │ ├── getTopic.ts # get_topic handler
│ │ ├── listTopics.ts # list_topics handler
│ │ ├── getContentGaps.ts # get_content_gaps handler
│ │ └── getBuildStatus.ts # get_build_status handler
│ ├── data/
│ │ ├── corpus.ts # Sample Actian documentation topics
│ │ ├── analyticsData.ts # Mock GA4 site search data
│ │ └── jenkinsData.ts # Mock Jenkins build history
│ └── lib/
│ └── search.ts # TF-IDF search engine
├── scripts/
│ └── ga4_export.sql # BigQuery query for real GA4 export
├── config/
│ ├── claude_desktop_config.json # Claude Desktop setup
│ └── mcp.json # Claude Code project setup
├── .env.example
└── tsconfig.json
```
---
## Tech stack
- **Runtime:** Node.js 18+ / TypeScript
- **MCP SDK:** `@modelcontextprotocol/sdk`
- **Search:** TF-IDF (built-in), upgradeable to pgvector / Chroma
- **Analytics:** Mock GA4 data, upgradeable to BigQuery
- **CI status:** Mock Jenkins data, upgradeable to live Jenkins REST API
---
## Related projects
- [`knowflow`](https://github.com/Bipin-24/knowflow) — MCP server with a RAGAS-style retrieval evaluation layer
- [`knowledge-graphs-for-ia`](https://github.com/Bipin-24/knowledge-graphs-for-ia) — graph builder with a typed edge model over a documentation corpus
- [`Documentation-AI-Assistant`](https://github.com/Bipin-24/Documentation-AI-Assistant) — RAG pipeline and chat UI over a documentation corpus
- [IA Playbook](https://information-architecture-playbook.vercel.app) — reference architecture for AI-readable documentation
---
## License
MIT. See [LICENSE](LICENSE).
---
## Author
**Bipin Pandey** — Principal Information Architect
[LinkedIn](https://www.linkedin.com/in/bipin-pandey24/) · [Portfolio](https://bipin-24.github.io)
TDQS
Scored across 5 tools
Each tool targets a distinct operation: semantic search, content retrieval by ID, listing metadata, analytics for gaps, and build pipeline status. There is no overlap or ambiguity between them.
All tool names consistently follow the verb_noun pattern in snake_case (search_docs, get_topic, list_topics, get_content_gaps, get_build_status). The naming is uniform and predictable.
Five tools is a well-scoped set for a documentation server, covering search, retrieval, listing, analytics, and CI/CD status without unnecessary bloat. Each tool earns its place.
The core documentation workflow is covered: discover topics via search/list, retrieve full content, and identify gaps. The build status tool provides complementary operational visibility, making the surface complete for its purpose.