Skip to main content
Glama
README.md
# Idea-To-Prod

**You give it an idea. It gives you back working, tested code.**

Idea-To-Prod is a multi-agent AI platform. You describe an application
idea in one sentence, and 6 AI agents work together โ€” one after another โ€”
to design it, write the code, test it, and (optionally) deploy it.

It is built as an [MCP](https://modelcontextprotocol.io/) server, so any
MCP-compatible AI assistant (Claude Desktop, GitHub Copilot, etc.) can use
it as a tool.

---

## ๐Ÿ“‹ Table of Contents

- [How it works](#-how-it-works)
- [The 6 agents](#-the-6-agents)
- [Which AI models are used](#-which-ai-models-are-used)
- [Setup](#๏ธ-setup)
- [How to run it](#๏ธ-how-to-run-it)
- [What you get back](#-what-you-get-back)
- [Project structure](#-project-structure)
- [Notes on a few design decisions](#-notes-on-a-few-design-decisions)

---

## ๐Ÿ”„ How it works

```mermaid
flowchart LR
    Idea(["๐Ÿ’ก Your idea"]) --> A1

    A1["Agent 1
    High-Level Design"] -->|saves doc| Drive1[("Google Drive")]
    A1 --> A2

    A2["Agent 2
    Detailed Design"] -->|saves doc| Drive2[("Google Drive")]
    A2 -->|creates tasks| Jira[("Jira")]
    A2 --> A3

    A3["Agent 3
    Write Code"] -->|pushes code| GH1[("GitHub")]
    A3 --> A4

    A4["Agent 4
    Write Tests"] -->|pushes tests| GH2[("GitHub")]
    A4 --> A5

    A5{"Agent 5
    Run Tests"}
    A5 -->|โŒ failed, try again| A4
    A5 -->|โœ… passed| A6

    A6["Agent 6
    Deploy (optional)"] --> Done(["๐ŸŽ‰ Done - code ready"])
```

Each agent does **one job** and then hands off to the next one. If the
tests fail (Agent 5), the flow goes back to Agent 4 to fix the tests โ€”
up to 3 times โ€” before giving up and reporting what went wrong.

---

## ๐Ÿค– The 6 agents

| # | Agent | What it does | Saves to |
|---|-------|---------------|----------|
| 1 | **High-Level Design** | Reads your idea and writes a short design document: what the app does, who it's for, what technologies to use. | Google Drive |
| 2 | **Detailed Design** | Takes the design and breaks it into concrete, buildable development tasks. | Google Drive + Jira |
| 3 | **Code Generation** | Reads the tasks and writes the actual application code. | GitHub (new repository) |
| 4 | **Unit Test Generation** | Reads the code and writes tests for it. Uses a **different AI model** than Agent 3, so it's a genuine second opinion, not the same model checking its own work. | GitHub (same repository) |
| 5 | **Test Execution** | Actually runs the tests. If they fail, sends the failure details back to Agent 4. If they pass, the pipeline is done. | โ€” |
| 6 | **Deploy** *(optional)* | Publishes the finished, tested app online. Only runs if you ask for it. | Hosting provider |

---

## ๐Ÿง  Which AI models are used

Every agent uses an AI model to do its job, but **not the same one for
everything** โ€” this project specifically requires Agent 3 (writing code)
and Agent 4 (writing tests) to use two *different* models, so Agent 4 is
a real second opinion, not an echo of Agent 3.

Right now every agent runs on **OpenAI**, with two different models
(`gpt-4o` for the heavier design/coding work, `gpt-4o-mini` for the
lighter tasks). Swapping any single agent to a different provider
(Gemini, Claude, etc.) is a one-line change โ€” see
[`src/idea_to_prod/config/models.py`](src/idea_to_prod/config/models.py).

---

## ๐Ÿงฉ How the pieces connect (MCP)

This project speaks [MCP](https://modelcontextprotocol.io/) in **two
directions**:

- **As a server** โ€” it exposes exactly one tool, `ideaToProd(idea)`. This
  is what an AI assistant like Claude Desktop calls.
- **As a client** โ€” internally, each agent connects out to *other* MCP
  servers (Google Drive, Jira, GitHub, Playwright) to actually save
  documents, create tasks, push code, and run tests.

```
   You / Claude Desktop
          โ”‚
          โ”‚  calls  ideaToProd("build me a calculator app")
          โ–ผ
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚  Idea-To-Prod      โ”‚   โ† this project
  โ”‚  MCP Server        โ”‚
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
            โ”‚  the 6 agents call out to:
            โ–ผ
  Google Drive ยท Jira ยท GitHub ยท Playwright   โ† real services (or local mocks)
```

For testing without any real accounts, every one of those four services
has a **local mock** (`tests/mocks/`) that behaves like the real thing โ€”
the GitHub mock creates a real local git repository, and the Playwright
mock actually runs the generated tests with `pytest`. Each service can be
switched from mock to real independently, one at a time, using its own
`USE_MOCK_<SERVICE>` flag in `.env` (all default to `true`, meaning
mocked).

---

## โš™๏ธ Setup

**You'll need:**

- Python 3.11 or newer
- [`uv`](https://docs.astral.sh/uv/) (Python package manager)
- `git` (the GitHub mock uses it directly)
- Node.js (only needed once you connect real, non-mock services โ€” they
  run via `npx`)
- An [OpenAI API key](https://platform.openai.com/api-keys)

**Install:**

```bash
uv sync
cp .env.example .env
```

Then open `.env` and set `OPENAI_API_KEY` to your real key. Everything
else can stay as-is (mocked) for your first run.

---

## โ–ถ๏ธ How to run it

There are three ways to use it โ€” pick whichever fits what you're doing.

### 1. Smoke test โ€” fastest way to see it work

```bash
uv run pytest tests/test_smoke.py -s
```

Runs all 6 agents against the local mocks with one sample idea, and
prints each agent's progress as it happens. This makes real OpenAI API
calls (small cost).

### 2. Our own CLI client โ€” the interactive way

```bash
uv run idea-to-prod-client
```

Asks you for an idea, then runs the whole pipeline and shows live
progress in your terminal.

### 3. Claude Desktop โ€” the "real" MCP way

1. Copy [`claude_desktop_config.example.json`](claude_desktop_config.example.json)
   into Claude Desktop's MCP settings, filling in this project's folder
   path and your API key.
2. Restart Claude Desktop.
3. Ask it something like: *"Use ideaToProd to build a CLI todo list app."*

---

## ๐Ÿ“ฆ What you get back

**If the tests pass:** the generated application files, the generated
test files, how many retries it needed, and links to the design documents
and Jira tasks created along the way.

**If the tests keep failing** (after 3 retries): the best attempt it
made, plus a clear report explaining what's still broken โ€” instead of
hanging forever or silently returning broken code.

---

## ๐Ÿ“ Project structure

```
idea-to-prod/
โ”œโ”€โ”€ pyproject.toml
โ”œโ”€โ”€ .env.example
โ”œโ”€โ”€ claude_desktop_config.example.json
โ”œโ”€โ”€ src/idea_to_prod/
โ”‚   โ”œโ”€โ”€ config/          # settings + which AI model each agent uses
โ”‚   โ”œโ”€โ”€ tools/            # connects each agent to its MCP service
โ”‚   โ”œโ”€โ”€ agents/            # the 6 agents
โ”‚   โ”œโ”€โ”€ flow.py            # ties all 6 agents together, including the retry loop
โ”‚   โ”œโ”€โ”€ server.py          # the MCP server (exposes ideaToProd)
โ”‚   โ””โ”€โ”€ client.py          # a simple CLI client for trying it out
โ””โ”€โ”€ tests/
    โ”œโ”€โ”€ mocks/              # local stand-ins for Drive/Jira/GitHub/Playwright
    โ””โ”€โ”€ test_smoke.py       # end-to-end test
```

---

## ๐Ÿ“ Notes on a few design decisions

A few choices here aren't obvious, so they're written down:

- **The retry loop (Agent 5 โ†’ Agent 4) is a plain Python loop**, not a
  CrewAI "Flow" cycle. Two attempts at building it as a native Flow cycle
  didn't reliably repeat on a second try during testing, so it was
  rebuilt as a simple, predictable `while` loop instead. Details in
  [`flow.py`](src/idea_to_prod/flow.py).
- **The GitHub repository name is decided by code, not by the AI.** Early
  testing showed the AI could invent a repository name in its final
  summary that didn't match the one it actually used โ€” a classic AI
  "hallucination" that broke every step after it. Now the name is
  computed once, in plain code, and passed to every agent that needs it.
- **Real MCP servers don't all use the same tool names.** The tool names
  this project calls (e.g. `create_document`) are its own internal
  agreement, matched exactly by the local mocks. Connecting a real
  service may need a one-line name adjustment โ€” see the note at the top
  of [`tools/mcp_connection.py`](src/idea_to_prod/tools/mcp_connection.py).

---

## ๐Ÿ“„ License

MIT โ€” see [`LICENSE`](LICENSE).