StoryForge MCP Server
README.md
# StoryForge — AI Backlog & Requirements Assistant
Turn a plain-English business brief into a structured, Agile-ready product
backlog (Epics → User Stories → Acceptance Criteria) using an LLM — then
measure it, ground it in real documents, and act on it with an agent.
Built as a hands-on tour of the modern Gen-AI product stack: **LLM APIs,
structured output, prompt engineering, evals, RAG, agents/tools, and MCP.**
## System map
```mermaid
flowchart LR
brief([Plain-English brief]) --> gen
docs([Company BRD / PRD]) --> rag
subgraph THREE [3 · Ground]
rag[Retrieve relevant chunks]
end
subgraph ONE [1 · Structure]
gen[Generate backlog]
end
rag -. context .-> gen
gen --> backlog[(Structured backlog)]
subgraph TWO [2 · Measure]
eval[Score quality]
end
backlog --> eval
subgraph FOUR [4 · Act]
agent[Agent + tools] --> jira[(Jira board)]
end
backlog --> agent
agent -. exposed via .-> mcp{{MCP server}}
mcp -. any AI client .-> ext([Claude Desktop])
```
A brief and the company's documents flow in; the app **structures** a backlog,
**measures** its quality, **grounds** it in real documents, and **acts** on it
with an agent — reusable by any AI through MCP.
---
## What it demonstrates
Each phase adds one core Gen-AI concept, and each was built to be understood by
running it. Full write-ups (with plain-English explanations) live in
[`CONCEPTS.md`](CONCEPTS.md).
| Phase | Concept | What it does |
|------:|---------|--------------|
| **1** | Structured output, system prompts, output validation | Brief → typed `Backlog` object (Epics → Stories → Acceptance Criteria), forced into a Pydantic schema |
| **2** | Evals, golden datasets, regression testing | Rule-based scorers grade backlog quality against a fixed dataset, so prompt changes can be measured, not guessed |
| **3** | RAG, embeddings, vector search, grounding | Retrieves relevant chunks of the company's own documents to ground generation — with a relevance guardrail that refuses out-of-scope questions |
| **4** | Tool calling, agent loops, MCP | An agent pushes backlog stories to a (local) Jira board via tool calling, exposed over MCP for any AI client to use |
## Architecture
The app never talks to a specific LLM vendor directly. It talks to an
`LLMProvider` interface ([`backend/llm/base.py`](backend/llm/base.py)), and a
factory picks the concrete provider from the `LLM_PROVIDER` env var. Today that's
Gemini; adding Claude or OpenAI means writing one new class — nothing else
changes. The same seam carries generation, embeddings, and agent tool-calling.
```
backend/
core/ Phase 1 — schema (models.py), prompt (story_generator.py), CLI
evals/ Phase 2 — golden dataset, scorers, runner
rag/ Phase 3 — chunking, vector store, retriever, Q&A guardrail
agent/ Phase 4 — tools (local Jira), agent loop, MCP server
llm/ provider interface + Gemini implementation
```
---
## Setup
```bash
# 1. Create and activate a virtual environment
python -m venv venv
venv\Scripts\activate # Windows PowerShell
# source venv/bin/activate # macOS/Linux
# 2. Install dependencies
python -m pip install -r requirements.txt
# 3. Add your API key
copy .env.example .env # then edit .env and paste your Gemini key
```
Get a free Gemini API key at https://aistudio.google.com/apikey
`GEMINI_MODEL` accepts a **comma-separated fallback chain** — the app tries each
model in order and fails over if one is overloaded (see `.env.example`).
---
## Run it, phase by phase
```bash
# Phase 1 — brief → structured backlog
python -m backend.core.cli
python -m backend.core.cli "Build an app for tracking gym workouts"
# Phase 2 — evals
python -m backend.evals.test_scorers # offline: scorers vs a bad backlog (no API)
python -m backend.evals.runner myrun # live: generate + score the golden dataset
# Phase 3 — RAG
python -m backend.rag.demo # index a doc → retrieve → grounded backlog
python -m backend.rag.ask "How long before my appointment can I cancel?"
# Phase 4 — agent + tools
python -m backend.agent.tools.jira_demo # the tool alone (no AI)
python -m backend.agent.demo # the agent drives the tool from an instruction
```
### Connect the MCP server to Claude Desktop
Wrap the Jira tools as an MCP server so any MCP client can use them. Add this to
your `claude_desktop_config.json`, then fully restart Claude Desktop:
```json
{
"mcpServers": {
"storyforge-jira": {
"command": "C:\\path\\to\\storyforge\\venv\\Scripts\\python.exe",
"args": ["-m", "backend.agent.mcp_server"],
"cwd": "C:\\path\\to\\storyforge",
"env": { "PYTHONPATH": "C:\\path\\to\\storyforge" }
}
}
}
```
Then ask Claude to "create a Jira ticket for X" and watch it land in
`jira_board.json`.
---
## Roadmap
- [x] **Phase 1** — Brief → structured backlog (LLM API, system prompts, structured output)
- [x] **Phase 2** — Prompt engineering + evals (measure & improve story quality)
- [x] **Phase 3** — RAG (ground stories in uploaded BRD/PRD documents) + guardrails
- [x] **Phase 4** — Agents + tools + MCP (push to Jira, agentic refinement)
- [ ] **Phase 5** — Web UI + multi-user (the demo shell over everything)
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues