FlowMCP
by MaxL963
README.md
# FlowMCP
An MCP server that turns a codebase into **business-language flow diagrams** an AI can read and edit.
Ask "how does login work?" and get this — not a list of files:
```mermaid
flowchart TD
n0["POST Login"]
n1["Authenticate User"]
n2{"The submitted password matches the stored hash?"}
n3(["Success"])
n4(["Failure"])
n5["Check Password"]
n6["Bcrypt Compare"]
n7["Find User By Email"]
n0 --> n1
n1 -.-> n5
n1 --> n2
n2 -->|yes| n3
n2 -->|no| n4
n5 --> n6
n5 -.-> n7
```
Then say *"insert an OTP step before the password check"* and the diagram changes — and stays changed when the code is re-indexed.
Or say *"open it on a canvas"* and edit it by hand in the browser, where every drag, rename and inserted step lands in the same durable layer:

---
## Why it exists
An AI assistant reasoning about a system reads raw source files. That is expensive, lossy, and gets worse as the repo grows. FlowMCP builds an architecture graph once, keeps it correct as code changes, and serves flows from it. The graph — not the file tree — becomes what the assistant reasons over.
**The server contains no language model.** It is a deterministic extractor, a graph store, and a renderer. The *calling* AI supplies the business language and the server persists it. That is not a cost optimisation: it removes an entire class of failure where two models disagree about what a node means. There is one interpreter of intent, and it is already in the conversation with you.
---
## Requirements
- **Node 24+** (developed on 26). TypeScript runs natively via type-stripping — there is no build step.
- Nothing else. SQLite is stdlib.
## Install
```bash
git clone <this repo>
cd FlowMCP
npm install # also copies the tree-sitter grammars into grammars/
```
Verify:
```bash
npm test # 121 tests
npx tsc --noEmit # typecheck
```
## Try it without an MCP client
One command: index a project and open the canvas on one flow.
```bash
npm run canvas -- /path/to/your-project "login"
npm run canvas -- /path/to/your-project "checkout" --depth 3
```
If the query matches nothing it prints the closest matches from the index instead of a stack trace, so the second attempt lands. Ctrl+C to stop. This is also the fastest way to work on the canvas page itself — it is re-read from disk on every request.
## Register it with your client
FlowMCP indexes **one project per server instance**. The project path is `argv[2]`, falling back to the working directory.
**Claude Code:**
```bash
claude mcp add flowmcp -- node /abs/path/to/FlowMCP/src/index.ts /abs/path/to/your-project
```
**Claude Desktop** — in `claude_desktop_config.json`:
```json
{
"mcpServers": {
"flowmcp": {
"command": "node",
"args": [
"/abs/path/to/FlowMCP/src/index.ts",
"/abs/path/to/your-project"
]
}
}
}
```
Use absolute paths. A wrong root silently indexes nothing.
The graph is written to `<your-project>/.flowmcp/graph.db`. Add `.flowmcp/` to that project's `.gitignore`.
---
# Tutorial
Everything below is what you'd actually say to the AI, followed by the tool it reaches for. You never call these by hand.
## 1. Index the codebase
> "Index this project."
```jsonc
analyze_codebase { } // or { "path": "...", "force": true }
```
```json
{
"filesScanned": 4, "filesParsed": 4, "filesSkipped": 0,
"nodes": 7, "edges": 5,
"unresolvedCalls": 7, "lowConfidenceEdges": 2,
"durationMs": 67
}
```
Incremental by content hash — re-running only re-parses files that changed. `force: true` rebuilds everything derived, and **never** touches your annotations or edits.
Upgrading FlowMCP re-parses on its own. A release that changes how code is read would otherwise reach only the files that happen to change afterwards, leaving the rest of your graph as the old version saw it — so the first index after an upgrade rebuilds the derived tables whether hashes match or not. You will see `filesParsed` equal to the whole tree once, and never need to remember `force`.
Read the last two numbers. They tell you how much of this graph is inference rather than fact.
## 2. Ask for a flow
> "Show me the login flow."
```jsonc
generate_diagram { "query": "login" }
```
Matching is full-text over code names, route paths, **and any business titles you've written** — so once a node is titled *Authenticate User*, asking for "authentication" finds it even though the code says `login`.
### Reading the diagram
| shape | meaning |
|---|---|
| `["Step"]` | a step |
| `{"Condition?"}` | a decision |
| `(["Outcome"])` | a terminal — where the flow ends |
| `-->` solid | the code definitely does this |
| `-.->` **dotted** | **a heuristic guess** |
**Dotted edges matter.** Without a type checker, calls through dependency injection (`this.authService.verify()`) are resolved by naming convention. That is usually right and sometimes wrong, and the diagram says which is which rather than presenting a guess as fact.
## 3. Open it on a canvas instead of reading diagram source
A twenty-step flow pasted into a terminal is unreadable, and that is the wrong place to edit one anyway.
> "Open the login flow on a canvas."
```jsonc
open_canvas { "flow": "login" } // or { "flow": "login", "depth": 3, "open": false }
```
```json
{
"url": "http://127.0.0.1:60515/?t=f4978cd4…&q=login",
"flow": "login", "nodes": 8, "edges": 7, "unnamed": 6
}
```
It opens in your browser and draws the same flow: steps as boxes, decisions as diamonds, outcomes as pills, heuristic calls as dashed arrows.
| tool | op it writes |
|---|---|
| drag a box (shift-click for several) | `set_layout` |
| double-click, or **Rename** | `rename_node` |
| the description field in the side panel | `describe_node` |
| **+ Before** / **+ After** | `insert_node` |
| **Connect** — click source, then target | `connect` |
| click a line, **Disconnect** | `disconnect` |
| **Delete** | `delete_node` |
| **Merge** (2+ selected) / **Split** | `merge_nodes` / `split_node` |
Nothing is written while you work. Edits queue in the side panel, where you can drop any one of them, and **Save** appends the batch — to the same overlay `edit_diagram` writes to. So a box you dragged and a step you drew survive re-indexing, and the next `get_flow` reads them back without being told anything happened.
`unnamed` counts nodes still carrying a humanised code name. The canvas marks them amber, which makes the annotating pass in the next section a visible to-do list rather than a guess.
Two things to know:
- **The URL is the credential.** It binds `127.0.0.1` only and every route requires the token in it — any page in your browser can POST to localhost, and without that check one could rewrite your overlay. The token changes on every server restart.
- **A broad query is unreadable, and says so.** Past 40 steps the result carries a `warning`: narrow the query or drop `depth` to 2 or 3. Ranks wider than five boxes wrap onto another row rather than stretching the picture past legibility.
- **Diamonds and pills are read-only.** They are computed from the code on every read (§6.1), so no edit can address them. The canvas says so rather than queueing an op that would vanish.
## 4. Give steps business names
This is the step that makes the tool worth using.
> "Rename these to business language: the controller is 'Authenticate User', the service method is 'Check Password'."
```jsonc
annotate_nodes {
"annotations": [
{ "node_key": "controller:src/controllers/auth.controller.ts:AuthController.login",
"title": "Authenticate User",
"description": "Takes the submitted email and password." },
{ "node_key": "service:src/services/auth.service.ts:AuthService.verify",
"title": "Check Password" }
]
}
```
Titles are **durable**. They survive every future re-index, because they live in a layer that re-indexing never rebuilds. Write them once.
Unannotated nodes fall back to a humanised code name (`findByEmail` → *Find By Email*), and every node reports whether it was annotated or humanised, so the AI can see what still needs naming.
When two humanised names would read the same, the second gets the smallest suffix that separates them — the layer (*Book (page)* beside *Book (function)*), the file, or whatever part of the directory differs (*GET (rooms/[code])*). A title you wrote is never rewritten, even when it collides: a duplicate you chose is a statement. Framework plumbing is left out of flows entirely — `useState` and `useParams` stay in the index for `diagnose`, but they are not steps.
## 5. Edit the flow in plain English
> "Insert a 'Send OTP' step before Check Password."
```jsonc
edit_diagram {
"note": "Insert Send OTP before Check Password",
"ops": [{
"op": "insert_node",
"payload": {
"position": "before",
"anchor": "service:src/services/auth.service.ts:AuthService.verify",
"node": { "title": "Send OTP", "type": "service" }
}
}]
}
```
Callers are rewired through the new step automatically. Ask for the diagram again and it's there.
**Nine operations** are available: `insert_node`, `delete_node`, `rename_node`, `describe_node`, `connect`, `disconnect`, `merge_nodes`, `split_node`, `set_layout`. Every one records the sentence that produced it in `note`, so the edit log reads as a history of intent.
## 6. Change the code, keep the diagram
> "The code changed — resync."
```jsonc
sync_architecture { }
```
Re-parses changed files, replays your edits on top, and reports **drift**:
| category | meaning |
|---|---|
| `orphaned_op` | an edit points at code that no longer exists |
| `unimplemented` | a step you drew that has no code behind it yet |
| `stale_annotation` | a named node's source changed since you named it |
| `ambiguous_rename` | one symbol vanished and a similar one appeared — possibly renamed |
**Nothing auto-resolves.** An orphaned edit is *kept*, not dropped — the code may come back, or you may want to re-anchor it. Resolution is your decision, expressed as new edits.
## 7. Explore
```jsonc
search_architecture { "query": "password", "kind": "service", "limit": 10 }
trace { "from": "<node_key>", "direction": "forward" } // or "backward"
get_node { "node_key": "..." }
get_architecture { "scope": "src/auth" }
```
`trace forward` is the execution path; `backward` finds everything that depends on a node — the "what breaks if I change this" question.
## 8. Audit the graph before trusting it
> "What's wrong with this architecture?"
```jsonc
diagnose { } // or { "checks": ["cycles", "unresolved_calls"] }
```
Six checks: `orphans`, `cycles`, `dead_flows`, `low_confidence`, `unresolved_calls`, `drift`.
Two of them are about the graph's own honesty:
- **`low_confidence`** — edges that exist but were guessed.
- **`unresolved_calls`** — calls the extractor *saw* and could not attribute at all. No edge was drawn, so the diagram is incomplete there. This is the difference between "this function calls nothing" and "we couldn't work out what it calls", and without this check the second one is invisible.
`diagnose` reports. It never fixes anything.
## 9. Export and round-trip
```jsonc
export_diagram { "flow": "login", "format": "markdown", "path": "/abs/path/login.md" }
```
Six formats: `mermaid`, `markdown` (diagram + a prose table), `json`, `plantuml`, `drawio`, `excalidraw`. The last two honour `set_layout` positions; Mermaid auto-layouts and ignores them.
You can also **edit the exported text by hand and feed it back**:
```jsonc
import_diagram { "content": "<edited mermaid>", "format": "mermaid" }
```
Renames, insertions and removed edges become overlay operations — never direct writes to the derived graph. Exports embed an invisible `%% flowmcp:node <id> <key>` block so nodes are matched by identity; retyped diagrams fall back to matching by title, and anything ambiguous becomes an insert rather than a guess.
---
## Tool reference
| tool | arguments |
|---|---|
| `analyze_codebase` | `path?` `force?` |
| `get_flow` | `query` `depth?` `save_as?` |
| `generate_diagram` | `query` `mode?` `format?` `depth?` |
| `open_canvas` | `flow` `depth?` `open?` |
| `annotate_nodes` | `annotations[]` — `{node_key, title, description?}` |
| `edit_diagram` | `ops[]` `note?` |
| `sync_architecture` | — |
| `search_architecture` | `query` `kind?` `limit?` |
| `trace` | `from` `direction` `depth?` |
| `get_node` | `node_key` |
| `get_architecture` | `scope?` |
| `diagnose` | `checks?` |
| `export_diagram` | `flow` `format` `path?` |
| `import_diagram` | `content` `format` `flow?` `note?` |
**Diagram modes:** `business` (default — titles only, no filenames, no syntax), `system`, `api`, `dependency`. `sequence`, `database` and `state` are declared but not implemented; they fail loudly rather than silently substituting.
---
## How it works
```
code ──scan/parse──▶ derived layer (disposable; rebuilt every index)
+
overlay log (append-only; never rebuilt)
▼
effective graph (what every tool reads)
```
The derived layer is a pure function of your source tree. The overlay is an ordered log of your edits. Everything you read is the two composed.
That is why `analyze_codebase` is always safe to re-run: **it cannot destroy your work, because your work does not live in the layer it rebuilds.**
Nodes are keyed `{kind}:{path}:{qualified_name}` — never line numbers, so an edit above a function doesn't detach anything anchored to it.
### Confidence ladder
| confidence | how the call was resolved |
|---|---|
| 1.0 | same file, `this.method()`, a relative import, `new Cls()`, or a package import |
| 0.7 | member call narrowed to one file by a relative import |
| 0.6 | `this.field.method()` — dependency injection, matched by naming convention |
| 0.5 | last resort: exactly one node repo-wide has that name |
| *no edge* | more than one candidate, or none — recorded as an unresolved call |
## Languages
TypeScript/JavaScript, Python, Go, Java, C#, PHP.
Frameworks detected from your manifests (`package.json`, `requirements.txt`, `pyproject.toml`, `go.mod`, `pom.xml`, `build.gradle`, `*.csproj`, `composer.json`): Express, Fastify, NestJS, Koa, Hono, FastAPI, Flask, Django, Gin, Echo, Chi, net/http, Spring, JAX-RS, ASP.NET Core, Laravel, Symfony, Slim.
Adding a language is one file in `src/packs/` plus one line in the registry.
## Known limits
Read these before trusting a diagram on a large codebase.
1. **Call resolution is heuristic.** Dynamic dispatch, DI containers and interface polymorphism produce gaps. Surfaced through edge confidence and `unresolved_calls` — never hidden.
2. **Interface-heavy C#/Java lose edges.** `IAuthService.Verify` alongside `AuthService.Verify` is ambiguous, so no edge is drawn.
3. **Node identity depends on names.** Renaming a symbol detaches edits targeting it; reported as `ambiguous_rename`, never auto-resolved, because auto-resolving a wrong guess silently corrupts the log.
4. **Some routing is detected but not read:** Django's `urlpatterns`, JAX-RS split `@GET`/`@Path`, and Spring class-level `@RequestMapping` prefixes.
5. **Import can't carry everything.** Edge kinds, descriptions, and merge/split aren't round-tripped; decision and terminal nodes are computed per-read and can't be addressed by an edit.
6. **Parsing is single-threaded.** Fine for normal repos; a worker pool is the upgrade path.
7. **Flow seeding is lexical.** A flow whose vocabulary appears nowhere in code or annotations needs annotating first, or an explicit seed.
## Development
```bash
npm test # node:test, no framework
npx tsc --noEmit # typecheck only; there is no build
```
```
src/db/ schema, connection, all SQL
src/scan/ file walking and hashing
src/extract/ tree-sitter extraction and call resolution
src/packs/ one file per language
src/graph/ flow computation, overlay replay, drift
src/render/ one file per output format
src/canvas/ the editable canvas: a local server and one page
src/tools/ one thin file per MCP tool
test/fixtures/ a small app per language, each with an EXPECTED.md
```
Architecture rationale lives in `docs/superpowers/specs/2026-07-29-flowmcp-design.md`.
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues