Skip to main content
Glama
README.md
# UltraPopper

**Adversarial, Karl Popper–style review for Claude Code.** UltraPopper front-loads *refutation* into
software work: before a line of code is written, it forces a change through a `conjecture → adversarial
refute → refine → survive → implement` loop, so the design that ships is the one that **survived attack** —
not the first one that looked plausible.

Its distinctive value is one move: **it surfaces the emergent, interaction-level bugs you would never think
to write a test for.** In a blind head-to-head against a rigorous test-first methodology building the same
cache library, the test-first build wrote *more* tests (92 vs 78) yet still shipped a real correctness bug
(an expired-but-unswept entry counting toward LRU capacity); UltraPopper caught it during refutation —
*before* implementation — because "what breaks when TTL and LRU act on the same structure?" is its native
question. You cannot test a bug you never conceived; refutation is the step that conceives it.

UltraPopper ships as two pieces:
- a **model-free MCP server** (`popper`) — holds no model, does no inference; every tool returns a
  prompt/plan for Claude Code and persists artifacts under `.popper/`;
- a **Claude Code plugin** (`ultrapopper`) — a `using-ultrapopper` entrypoint (auto-loaded via a SessionStart
  hook) plus per-command skills that drive the server.

## Quickstart

Prereqs: **Node ≥ 18** and the **Claude Code CLI**.

```bash
# from a checkout of this repo — installs the server + plugin (user scope):
bash scripts/install.sh
# fresh machine (private repo): GH_IBM_TOKEN=github_pat_xxx bash scripts/install.sh
```

Then **restart Claude Code**, `cd` into any project, and ask it to build something non-trivial — the
refutation loop runs automatically before code is written. Or drive it explicitly:
`ultrapopper:understand → :clarify → :hypothesize → :refute → :implement` (`:solve` picks the next step).
Sanity check: `npm run verify && npm run smoke`.

---

## When to use it

UltraPopper is **insurance against the bug you'd never think to write a test for** — worth its cost only when
a mistake would be expensive *and* could arise from how pieces **interact**. It earns a real *correctness*
edge on **emergent, interaction-level** failures, and little on routine work the model already handles. (A
two-task benchmark, same model across arms: on a self-contained rate-limiter it added nothing — every arm was
correct; on a TTL×LRU cache it was the **only** approach that shipped correct code, because the interaction
bug is one a plain build "tests around" without ever conceiving.)

**Reach for it (light → ultra) when:**
- You're implementing an idea into a **large / unfamiliar codebase** — the risk is interaction with code you
  can't hold in your head (shared state, ordering, invariants). Refutation grounds in the *actual* code
  (`understand` + evidence anchors), catching what the model can't conceive on its own.
- **Multiple components interact** — caches, concurrency, TTLs, cross-module contracts.
- The blast radius is **money / prod / security / data**, or the change is hard to reverse.
- You're **extending prior work** in a recurring domain — the cross-session knowledgebase feeds last time's
  failure patterns into this refutation, and sharpens with reuse.
- The design space is **wide or ambiguous** and you want to settle it *before* writing code.

**Skip it (just answer, or just edit) when:** a rename / typo / constant tweak / one obvious localized
function; a self-contained thing the model clearly knows cold; a reversible, well-specified change; or a
question with nothing to build.

> **The one question:** *"Could this fail in a way I wouldn't test for — because of how it interacts with
> something else — and would that failure be expensive?"* If yes, refute **first** (it only works before you
> code). The plugin triages this automatically (skip / light / ultra); say *"just do it"* to force a skip.

## Why model-free?

The server enforces *process*, not *intelligence*. It never calls a model. Every tool is a two-call pattern:
call 1 returns a prompt/plan for Claude Code to reason over; call 2 accepts the reasoned artifact and persists
it. Claude does all the thinking; the server guarantees the discipline — a locked contract, a firewall, a
durable knowledgebase, and a refutation record. This makes the loop deterministic, auditable, and impossible
to shortcut.

## How it works

UltraPopper is two cooperating pieces: a Claude Code **plugin** (`ultrapopper`) and the model-free **MCP
server** (`popper`) it drives. Together they run an adversarial review *before* Claude writes code, so what
ships is a design that survived attack rather than the first one that looked plausible.

When you ask Claude to build something non-trivial, the flow is:

1. **Triage** — the plugin sizes up the request and picks a gear: **skip** (just answer or make the edit),
   **light** (one quick refutation pass), or **ultra** (the full loop) — so the review only runs when it is
   worth the cost.
2. **Understand** — it indexes your codebase once, so the review is grounded in the real code.
3. **Clarify** — it restates the task as a locked contract (assumptions, constraints, open questions) and
   settles anything ambiguous with you first.
4. **Hypothesize** — it proposes one or more candidate designs (several independent ones, from different
   angles, in ultra mode).
5. **Refute** — it then *attacks* those designs with isolated probes that try to break each piece and each
   interaction, and prunes whatever fails. This is the step that surfaces the emergent, interaction-level bug
   you would never think to write a test for. It loops (refine → refute) until a design survives.
6. **Implement** — finally it writes the code and tests for the design that survived.

The MCP server does no thinking of its own — it holds no model and makes no network calls. It enforces the
*process* (the locked contract, the isolation between proposing and attacking, a durable refutation record,
and a knowledgebase that carries lessons across sessions) while Claude does the actual reasoning. Because it
is a standard MCP server, the same engine can back any MCP-capable client, not just Claude Code.

## The loop

```
understand → clarify (lock a contract) → hypothesize → refute → [refine ↺] → implement
```

| Tool | What it does |
|---|---|
| `popper_understand` | Mechanically index the codebase into `.popper/`; second call attaches Claude's architecture/conventions/invariants. Returns a **size-bounded** summary (full profile persisted) so a huge repo can't blow the result limit. |
| `popper_clarify` | Turn a raw statement into a **locked problem contract**. Refuses to silently abandon an in-progress problem (pass `newProblem: true` to start fresh). |
| `popper_hypothesize` | Draft (call 1) then persist (call 2) a split **public/private** hypothesis. At `depth: "ultra"`, fans out K diverse Proposers and decomposes their solutions into a shared **fragment vocabulary**. |
| `popper_refute` | Return a **firewalled** fan-out plan of isolated adversarial probes; log verdicts and decide survive / refine / exhaust. |
| `popper_implement` | Emit a one-shot implementation prompt for the surviving hypothesis / best composite. |
| `popper_solve` | Report the current phase and the next tool to call (the conductor). |
| `popper_status` | Summarize the profile, phase, knowledgebase, and contract stats. |

### Light vs. ultra
- **`light`** — one hypothesis, one refutation fan-out. For a self-contained 1–2 file change.
- **`ultra`** — K diverse Proposers → fragment decomposition → combinatorial viable-path enumeration →
  per-fragment **and** interacting-pair **and** N-way composite refutation → best-composite assembly. For
  multi-file, architecturally significant work with several viable approaches.

## The information firewall (load-bearing)

A hypothesis is split into **public** (solution, assumptions) and **private** (reasoningTrace, confidence,
rejectedAlternatives) fields. Two gates in `src/firewall.ts` are the *sole* boundary to a Refuter:
- `toRefuterView(hypothesis)` — strips the private fields.
- `toCanonFragmentView(fragment)` — exposes only `{ key, files, description }` (never `proposers`).

Nothing else may hand data to a probe. The end-to-end smoke test proves it from both sides: the private
reasoning is retained in the private session log yet **absent** from every public artifact.

## Ultra internals

- **Fragments** (`src/fragments.ts`) — proposers decompose solutions into toggleable sub-changes with a shared
  kebab-case `key`, plus `requires`/`conflictsWith`. Keys are typographically canonicalized, and a two-call
  `mergePlan` lets Claude confirm semantic aliases (model-free — the server proposes candidates by file
  overlap, it never *decides* equivalence).
- **Combinatorial engine** (`src/bdd.ts`) — enumerates conflict-free, requires-closed **viable composites**
  with requires-satisfiability forward-pruning, reduced to the maximal frontier.
- **Interacting-pair + N-way composite probes** — beyond per-fragment probes, UltraPopper probes co-occurring
  **pairs** (capped/ranked by file overlap) and whole **composites** for *emergent* fatals that no fragment
  reveals alone. A grounded fatal pair adds a conflict edge; a fatal composite forbids that combination.
- **Cross-session knowledgebase** (`src/kb.ts`) — grounded fatal patterns and **resolved-defect** findings are
  promoted to `.popper/kb.json`; on later problems, patterns on overlapping files are fed to Refuters as
  ammunition (grounded by file-hash; stale ones trigger re-verification). The loop learns from its own work.
- **File-based intake** — large K-proposer payloads are written to `.popper/inbox/` and read by the server,
  never passed as one fragile inline JSON argument.

## The three deepenings

Recent work sharpened the *core* (refutation quality) rather than broadening scope:
- **Refutation report** (`.popper/reports/<id>.json`) — a durable, first-class record of what each round
  probed, pruned, suspected, and left uncovered. The loop's distinctive artifact.
- **Coverage critic** — the plan exposes the failure-mode taxonomy as `coverageTargets`; the verdict returns
  the `coverageGap` (classes no probe addressed), so blind spots are named, not assumed away.
- **Refuter calibration** — a fatal prunes only if the Refuter *demonstrated* it (`reproduced: true`); an
  unreproduced hunch is recorded as **suspected**, not acted on, so it can't kill a good fragment.

## Scratch auto-cleanup

When a contract's loop **finishes** — `popper_implement` completes, `popper_refute` exhausts, or the
contract is abandoned via `popper_clarify { newProblem: true }` — its transient scratch is purged
automatically so it can't accumulate and hog disk/context:

- `.popper/bdd-trees/<id>.json` — the viable-path tree; consumed by `implement`, never read across contracts.
- `.popper/inbox/*.json` — proposer staging files, already read into the session log.
- stray `*.tmp` — interrupted atomic-write intermediaries.

Durable artifacts are **never** touched: `reports/` (the SuperPoppers handoff), `solutions/`,
`refutations/` (session logs), `kb.json`, `knowledgebase.json`, `state.json`, `config.yaml`. The purge is
best-effort (it never fails a run) and defaults **on** — opt out per repo with `cleanup: { purge_scratch: false }`
in `config.yaml`. (Unlike `max_patterns`, this loses no real data, so it is safe to default on.)

## Privacy, cost & control

- **Privacy — nothing leaves your machine, nothing lands in git.** The server is model-free and does **no
  network I/O**; it only writes under `.popper/`. That dir holds the model's *private* reasoning traces
  (session logs) plus the KB/reports/trees, so on first run UltraPopper writes a self-contained
  `.popper/.gitignore` of `*` — git ignores the whole working dir (including that file), so nothing under
  `.popper/` is ever accidentally committed. Delete it or `git add -f` to deliberately commit (e.g. to
  share a KB across a team).
- **Scoped, gitignore-aware scanning.** `popper_understand` stops at nested git-repo boundaries, caps at
  `MAX_SCAN_FILES`, honors the repo's `.gitignore` on top of `DEFAULT_IGNORES`, and takes a `path` arg to
  scope to a single subdirectory — a huge or multi-repo working dir can't pollute the profile.
- **Cost & control.** The SessionStart entrypoint **triages every request (skip / light / ultra)** so the
  loop runs only when it earns its cost; **ultra is token-heavy** (K proposers + a probe per fragment and
  per interacting pair). Say *"skip the loop"* to bypass for one task, or
  `claude plugin uninstall ultrapopper@ultrapopper` to disable it entirely.

## Known limitations

- One active popper session per working dir — two concurrent sessions can race on `.popper/state.json`.
- `.gitignore` awareness covers the common subset (dir-only, anchored, `*`/`**`/`?`, negation); it does not
  read per-directory nested `.gitignore` files.

## The plugin (`ultrapopper`)

The repo is itself a local plugin **marketplace**. Installing the plugin registers a SessionStart hook that
auto-loads the `using-ultrapopper` entrypoint every session, plus per-command skills invoked as
`ultrapopper:understand | clarify | hypothesize | refute | implement | solve`. A plain *"build X"* request
then drives the whole loop automatically — refute before code, prove the firewall, then implement.

### Install

```bash
# from a checkout (server + plugin, user scope):
bash scripts/install.sh
```

This builds the server, registers the `popper` MCP server (`claude mcp add popper -s user`), and installs the
plugin (`claude plugin marketplace add . && claude plugin install ultrapopper@ultrapopper`). Restart Claude
Code so the SessionStart hook loads, then `cd` into any project and ask it to build something non-trivial.

## Quality & testing

The whole project is built and hardened under a strict, ratcheting gate:

```bash
npm test             # full vitest suite
npm run verify       # tsc build + tests + coverage, then RATCHET .baseline.json up on success
npm run verify:check # same gate, read-only (CI / pre-push)
npm run smoke        # end-to-end: spawn the built stdio server, drive the whole pipeline, prove the firewall
```

`scripts/verify.mjs` fails on any regression below the `.baseline.json` floor (test count + coverage) and
raises the floor on success. `.github/workflows/ci.yml` runs `npm ci → verify:check → smoke` on Node 18/20/22.
Current floor: **441 tests, ~92.5% branch coverage**, statements/lines ~99%. Every feature and every fixed bug
landed with a test in the same commit; every adversarial-review finding was verified before fixing.

**Prove it's an improvement, don't assert it.** `npm run compare` builds the archived baseline in a throwaway
worktree and diffs the *same inputs* through both versions' (model-free, deterministic) handlers — it doubles
as a cross-version regression gate (non-zero exit if the current version is worse on any probe). Latest run:
**6 improved, 0 regressed** — see [docs/OLD-VS-NEW.md](docs/OLD-VS-NEW.md).

## SuperPoppers

UltraPopper composes with test-first methodologies rather than replacing them: run the refutation loop to
catch what you'd never test for, then hand the **refutation report** to a TDD executor for broad routine
coverage. That pipeline — **[SuperPoppers](https://github.com/Pranaav003/superpoppers)** — measured
best-of-both in the bake-off: as correct as UltraPopper *and* better-tested than pure TDD.

## Layout

```
src/            # model-free server: handlers, firewall, bdd, kb, fragments, compression, report, …
plugins/        # the ultrapopper Claude Code plugin (skills + SessionStart hook)
scripts/        # install.sh, verify.mjs (ratchet gate), smoke.mjs (e2e)
.claude-plugin/ # marketplace manifest
docs/           # design spec + phase plans
```

## License

[MIT](LICENSE).

TDQS

A3.8/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a distinct phase in the workflow (understand, status, clarify, hypothesize, implement, refute, solve, learn), with no overlapping purposes. The descriptions clearly differentiate their roles.

Naming Consistency5/5

All tool names follow the consistent pattern 'popper_<verb>', using snake_case throughout. The naming is predictable and systematic.

Tool Count5/5

With 8 tools covering all major phases of the scientific problem-solving workflow, the count is well-scoped. Each tool earns its place without unnecessary bloat or omission.

Completeness5/5

The tool set provides complete coverage of the intended workflow: from codebase understanding to problem clarification, hypothesis, implementation, refutation, orchestration, and learning. No obvious gaps are present.

Maintenance

ActivityMaintained
ResponsivenessSyncing