Skip to main content
Glama
ankushchaudhari203-commits

StoryForge MCP Server

README.md
# StoryForge — AI Backlog & Requirements Assistant

Turn a plain-English business brief into a structured, Agile-ready product
backlog (Epics → User Stories → Acceptance Criteria) using an LLM — then
measure it, ground it in real documents, and act on it with an agent.

Built as a hands-on tour of the modern Gen-AI product stack: **LLM APIs,
structured output, prompt engineering, evals, RAG, agents/tools, and MCP.**

## System map

```mermaid
flowchart LR
  brief([Plain-English brief]) --> gen
  docs([Company BRD / PRD]) --> rag
  subgraph THREE [3 · Ground]
    rag[Retrieve relevant chunks]
  end
  subgraph ONE [1 · Structure]
    gen[Generate backlog]
  end
  rag -. context .-> gen
  gen --> backlog[(Structured backlog)]
  subgraph TWO [2 · Measure]
    eval[Score quality]
  end
  backlog --> eval
  subgraph FOUR [4 · Act]
    agent[Agent + tools] --> jira[(Jira board)]
  end
  backlog --> agent
  agent -. exposed via .-> mcp{{MCP server}}
  mcp -. any AI client .-> ext([Claude Desktop])
```

A brief and the company's documents flow in; the app **structures** a backlog,
**measures** its quality, **grounds** it in real documents, and **acts** on it
with an agent — reusable by any AI through MCP.

---

## What it demonstrates

Each phase adds one core Gen-AI concept, and each was built to be understood by
running it. Full write-ups (with plain-English explanations) live in
[`CONCEPTS.md`](CONCEPTS.md).

| Phase | Concept | What it does |
|------:|---------|--------------|
| **1** | Structured output, system prompts, output validation | Brief → typed `Backlog` object (Epics → Stories → Acceptance Criteria), forced into a Pydantic schema |
| **2** | Evals, golden datasets, regression testing | Rule-based scorers grade backlog quality against a fixed dataset, so prompt changes can be measured, not guessed |
| **3** | RAG, embeddings, vector search, grounding | Retrieves relevant chunks of the company's own documents to ground generation — with a relevance guardrail that refuses out-of-scope questions |
| **4** | Tool calling, agent loops, MCP | An agent pushes backlog stories to a (local) Jira board via tool calling, exposed over MCP for any AI client to use |

## Architecture

The app never talks to a specific LLM vendor directly. It talks to an
`LLMProvider` interface ([`backend/llm/base.py`](backend/llm/base.py)), and a
factory picks the concrete provider from the `LLM_PROVIDER` env var. Today that's
Gemini; adding Claude or OpenAI means writing one new class — nothing else
changes. The same seam carries generation, embeddings, and agent tool-calling.

```
backend/
  core/     Phase 1 — schema (models.py), prompt (story_generator.py), CLI
  evals/    Phase 2 — golden dataset, scorers, runner
  rag/      Phase 3 — chunking, vector store, retriever, Q&A guardrail
  agent/    Phase 4 — tools (local Jira), agent loop, MCP server
  llm/      provider interface + Gemini implementation
```

---

## Setup

```bash
# 1. Create and activate a virtual environment
python -m venv venv
venv\Scripts\activate        # Windows PowerShell
# source venv/bin/activate   # macOS/Linux

# 2. Install dependencies
python -m pip install -r requirements.txt

# 3. Add your API key
copy .env.example .env       # then edit .env and paste your Gemini key
```

Get a free Gemini API key at https://aistudio.google.com/apikey

`GEMINI_MODEL` accepts a **comma-separated fallback chain** — the app tries each
model in order and fails over if one is overloaded (see `.env.example`).

---

## Run it, phase by phase

```bash
# Phase 1 — brief → structured backlog
python -m backend.core.cli
python -m backend.core.cli "Build an app for tracking gym workouts"

# Phase 2 — evals
python -m backend.evals.test_scorers      # offline: scorers vs a bad backlog (no API)
python -m backend.evals.runner myrun      # live: generate + score the golden dataset

# Phase 3 — RAG
python -m backend.rag.demo                # index a doc → retrieve → grounded backlog
python -m backend.rag.ask "How long before my appointment can I cancel?"

# Phase 4 — agent + tools
python -m backend.agent.tools.jira_demo   # the tool alone (no AI)
python -m backend.agent.demo              # the agent drives the tool from an instruction
```

### Connect the MCP server to Claude Desktop

Wrap the Jira tools as an MCP server so any MCP client can use them. Add this to
your `claude_desktop_config.json`, then fully restart Claude Desktop:

```json
{
  "mcpServers": {
    "storyforge-jira": {
      "command": "C:\\path\\to\\storyforge\\venv\\Scripts\\python.exe",
      "args": ["-m", "backend.agent.mcp_server"],
      "cwd": "C:\\path\\to\\storyforge",
      "env": { "PYTHONPATH": "C:\\path\\to\\storyforge" }
    }
  }
}
```

Then ask Claude to "create a Jira ticket for X" and watch it land in
`jira_board.json`.

---

## Roadmap

- [x] **Phase 1** — Brief → structured backlog (LLM API, system prompts, structured output)
- [x] **Phase 2** — Prompt engineering + evals (measure & improve story quality)
- [x] **Phase 3** — RAG (ground stories in uploaded BRD/PRD documents) + guardrails
- [x] **Phase 4** — Agents + tools + MCP (push to Jira, agentic refinement)
- [ ] **Phase 5** — Web UI + multi-user (the demo shell over everything)