Skip to main content
Glama
Bipin-24

actian-docs-mcp

by Bipin-24
README.md
# docs-mcp

![TypeScript](https://img.shields.io/badge/TypeScript-5.x-3178C6?logo=typescript&logoColor=white)
![Node](https://img.shields.io/badge/Node-18%2B-339933?logo=node.js&logoColor=white)
![MCP](https://img.shields.io/badge/MCP-Model_Context_Protocol-D97757)
![License](https://img.shields.io/badge/License-MIT-blue)

An [MCP (Model Context Protocol)](https://modelcontextprotocol.io) server that exposes Actian product documentation, GA4 content analytics, and Jenkins CI build status to Claude, enabling natural-language queries across all three sources in a single conversation.

> **Runs with no credentials.** The corpus, GA4 analytics, and Jenkins data ship as realistic mock fixtures, so you can clone, build, and query it in about two minutes. Every source has a documented upgrade path to live data. See [Live data](#live-data).

```
Claude Desktop / Claude Code
        │
        │  JSON-RPC over stdio
        ▼
  docs-mcp
        │
        ├── search_docs      → Actian Markdown corpus (TF-IDF, swap for pgvector)
        ├── get_topic        → Full topic content by ID
        ├── list_topics      → Corpus index with product/type filters
        ├── get_content_gaps → GA4 zero-result search queries (mock → BigQuery)
        └── get_build_status → Jenkins publish pipeline status (mock → live API)
```

## What this enables

Once connected, you can ask Claude:

> *"What are the top 10 search queries on the Actian docs site that returned no results this month?"*

> *"Find every topic in the Analytics Engine 8.0 corpus that mentions the Tableau connector, and check whether the publish pipeline is currently green."*

> *"Did the docs build pass? If it failed, search the corpus for the RPM upgrade procedure and show me the relevant steps."*

## Demo

<!-- TODO: record a short GIF of Claude calling get_content_gaps and returning
     a ranked list, save it as docs/demo.gif, then uncomment the line below.
     This is the single highest-value addition to this README. -->

<!-- ![Claude calling get_content_gaps](docs/demo.gif) -->

---

## Quick start

### 1. Clone and install

```bash
git clone https://github.com/Bipin-24/docs-mcp.git
cd docs-mcp
npm install
npm run build
```

### 2. Connect to Claude Desktop

Open your Claude Desktop config file:

| OS | Path |
|----|------|
| macOS | `~/Library/Application Support/Claude/claude_desktop_config.json` |
| Windows | `%APPDATA%\Claude\claude_desktop_config.json` |

Add this block (replace the path with your actual clone location):

```json
{
  "mcpServers": {
    "actian-docs": {
      "command": "node",
      "args": ["/absolute/path/to/docs-mcp/dist/index.js"]
    }
  }
}
```

Restart Claude Desktop. You should see `actian-docs` in the tools list.

### 3. Connect to Claude Code

Drop a `.mcp.json` file in your project root:

```json
{
  "mcpServers": {
    "actian-docs": {
      "command": "node",
      "args": ["../docs-mcp/dist/index.js"]
    }
  }
}
```

---

## Tools

### `search_docs`

Keyword search across the documentation corpus, ranked by TF-IDF with title and tag match bonuses. Upgradeable to embedding-based semantic search. See [Architecture](#search).

```
Input:
  query    string  required  Natural language search query
  product  string  optional  analytics-engine | ingres | actian-client | all
  version  string  optional  e.g. "8.0", "11.x"
  limit    number  optional  1–10, default 5

Output: ranked list of matching topics with excerpt and relevance score
```

**Example prompt:** *"How do I upgrade Analytics Engine using RPM packages?"*

---

### `get_topic`

Retrieve the full Markdown content of a topic by ID.

```
Input:
  topic_id  string  required  Topic ID from search_docs results

Output: full topic content with metadata
```

---

### `list_topics`

Browse the corpus index.

```
Input:
  product     string  optional  analytics-engine | ingres | actian-client | all
  topic_type  string  optional  concept | task | reference | troubleshooting | all

Output: topics grouped by product
```

---

### `get_content_gaps`

Surfaces search queries that returned zero or very few results, meaning documentation your users need but that does not exist yet.

```
Input:
  days          number  optional  Lookback window, default 30
  limit         number  optional  Max gaps to return, default 20
  min_searches  number  optional  Minimum search volume, default 2

Output: ranked gap list with gap type, search volume, and nearest existing topics
```

**Example prompt:** *"What are the top missing documentation topics based on search data from the last 30 days?"*

---

### `get_build_status`

Returns the current Jenkins CI pipeline status for the docs publish job.

```
Input:
  job  string  optional  Jenkins job name, default "actian-docs-publish"

Output: latest build result, stage breakdown, and recent history
```

---

## Architecture

### Search

The server ships with a lightweight TF-IDF keyword search engine (`src/lib/search.ts`) that requires no external dependencies or API keys. It scores documents using term frequency with title and tag match bonuses.

To upgrade to embedding-based semantic search:

1. Add `chromadb` to `package.json`
2. Run `npm run index` to embed the corpus using the Python indexer (`scripts/index_corpus.py`)
3. Swap `scoreTopics()` in `src/lib/search.ts` for a Chroma similarity query

### Live data

The server ships with realistic mock data for GA4 analytics and Jenkins. To switch to live sources, add credentials to `.env`:

```bash
cp .env.example .env
# Edit .env with your BigQuery project ID and Jenkins token
```

See `.env.example` for all available configuration options.

The BigQuery query for GA4 Site Search export is in `scripts/ga4_export.sql`.

---

## Project structure

```
docs-mcp/
├── src/
│   ├── index.ts              # MCP server — tool registry and routing
│   ├── tools/
│   │   ├── searchDocs.ts     # search_docs handler
│   │   ├── getTopic.ts       # get_topic handler
│   │   ├── listTopics.ts     # list_topics handler
│   │   ├── getContentGaps.ts # get_content_gaps handler
│   │   └── getBuildStatus.ts # get_build_status handler
│   ├── data/
│   │   ├── corpus.ts         # Sample Actian documentation topics
│   │   ├── analyticsData.ts  # Mock GA4 site search data
│   │   └── jenkinsData.ts    # Mock Jenkins build history
│   └── lib/
│       └── search.ts         # TF-IDF search engine
├── scripts/
│   └── ga4_export.sql        # BigQuery query for real GA4 export
├── config/
│   ├── claude_desktop_config.json  # Claude Desktop setup
│   └── mcp.json                    # Claude Code project setup
├── .env.example
└── tsconfig.json
```

---

## Tech stack

- **Runtime:** Node.js 18+ / TypeScript
- **MCP SDK:** `@modelcontextprotocol/sdk`
- **Search:** TF-IDF (built-in), upgradeable to pgvector / Chroma
- **Analytics:** Mock GA4 data, upgradeable to BigQuery
- **CI status:** Mock Jenkins data, upgradeable to live Jenkins REST API

---

## Related projects

- [`knowflow`](https://github.com/Bipin-24/knowflow) — MCP server with a RAGAS-style retrieval evaluation layer
- [`knowledge-graphs-for-ia`](https://github.com/Bipin-24/knowledge-graphs-for-ia) — graph builder with a typed edge model over a documentation corpus
- [`Documentation-AI-Assistant`](https://github.com/Bipin-24/Documentation-AI-Assistant) — RAG pipeline and chat UI over a documentation corpus
- [IA Playbook](https://information-architecture-playbook.vercel.app) — reference architecture for AI-readable documentation

---

## License

MIT. See [LICENSE](LICENSE).

---

## Author

**Bipin Pandey** — Principal Information Architect  
[LinkedIn](https://www.linkedin.com/in/bipin-pandey24/) · [Portfolio](https://bipin-24.github.io)

TDQS

A4.3/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct operation: semantic search, content retrieval by ID, listing metadata, analytics for gaps, and build pipeline status. There is no overlap or ambiguity between them.

Naming Consistency5/5

All tool names consistently follow the verb_noun pattern in snake_case (search_docs, get_topic, list_topics, get_content_gaps, get_build_status). The naming is uniform and predictable.

Tool Count5/5

Five tools is a well-scoped set for a documentation server, covering search, retrieval, listing, analytics, and CI/CD status without unnecessary bloat. Each tool earns its place.

Completeness5/5

The core documentation workflow is covered: discover topics via search/list, retrieve full content, and identify gaps. The build status tool provides complementary operational visibility, making the surface complete for its purpose.

Maintenance

ActivitySlowing
ResponsivenessNo issues