Skip to main content
Glama
donaldrichard19-LVD

ui-component-judgment-mcp

README.md
# Pattern

[![Publish](https://github.com/donaldrichard19-LVD/pattern-mcp/actions/workflows/publish.yml/badge.svg)](https://github.com/donaldrichard19-LVD/pattern-mcp/actions/workflows/publish.yml)
[![npm version](https://img.shields.io/npm/v/pattern-mcp.svg)](https://www.npmjs.com/package/pattern-mcp)
[![npm downloads](https://img.shields.io/npm/dt/pattern-mcp.svg)](https://www.npmjs.com/package/pattern-mcp)
[![MIT license](https://img.shields.io/badge/license-MIT-111111.svg)](./LICENSE)

Pattern checks a coding agent's UI decisions against your own design
system before it builds, whether that lives in code or in Figma. It helps
the agent reuse what the product already has and build only what is
actually missing. An opt-in
enforcement boundary can make that check required, and every decision
is recorded in an auditable ledger you can verify later.

[Website](https://usepattern.sh) · [npm](https://www.npmjs.com/package/pattern-mcp) · [Report an issue](https://github.com/donaldrichard19-LVD/pattern-mcp/issues/new/choose)

<details>
<summary><strong>Contents</strong> (click to expand)</summary>

- [Install](#install) · [What Pattern Does](#what-pattern-does) · [How it works](#how-it-works) · [Quick Start](#quick-start) · [Recommended workflow](#recommended-workflow) · [Try it](#try-it) · [Validation examples](#validation-examples)
- [Enforcement boundary: hook + CI gate](#enforcement-boundary-hook--ci-gate) (require the call, don't just log it)
- **Core tools** (on by default -- see [Tool tiers](#tool-tiers)): [`register_design_system`](#tool-register_design_system) · [`recommend_component`](#tool-recommend_component) · [`extract_requirements`](#tool-extract_requirements) · [`verify_component`](#tool-verify_component) · [`record_component_decision`](#tool-record_component_decision)
- **[Advanced tools](#advanced-tools)** (`PATTERN_TOOLS=full`): [`get_figma_evidence`](#tool-get_figma_evidence) · [`read_ledger`](#tool-read_ledger) · [`report_build_cost`](#tool-report_build_cost) · [`report_outcome_proxy`](#tool-report_outcome_proxy) · [Feature cost attribution](#feature-cost-attribution) · [Outcome proxies](#outcome-proxies) · [Per-project judgment ledger](#per-project-judgment-ledger) · [`check_ledger_liveness`](#tool-check_ledger_liveness) · [`sweep_ledger_liveness`](#tool-sweep_ledger_liveness) · [`export_ledger_provenance`](#tool-export_ledger_provenance) · [`backfill_ledger_snapshot_ref`](#tool-backfill_ledger_snapshot_ref) · [`post_ledger_provenance_to_github`](#tool-post_ledger_provenance_to_github) · [Ledger integrity and decision provenance](#ledger-integrity-and-decision-provenance)
- [Per-project decision memory](#per-project-decision-memory) · [Security and privacy](#security-and-privacy) · [Telemetry](#telemetry)
- **Cost:** [The `_meta` field](#the-_meta-field) · [Prompt caching](#prompt-caching) · [Measured cache and fetch behavior](#measured-cache-and-fetch-behavior-historical-pre-018) · [Search limits](#search-limits-removed-in-018) · [Ensemble cost](#ensemble-cost-boundary-risk-cases-only) · [Session call cap](#session-call-cap)
- [Local call log](#local-call-log) · [Known limitations](#known-limitations)

</details>

## Install

```bash
npx pattern-mcp init
```

This is the only command you run yourself. It downloads Pattern,
detects which MCP client(s) you have (Claude Code, Claude Desktop,
Cursor, or Codex CLI), connects each one for you, and offers to add
your Anthropic API key. Then register your design system (see
[`register_design_system`](#tool-register_design_system)) -- Pattern has
nothing to judge against until you do. See [Quick Start](#quick-start) below for what it
does step by step, or
[Connect Pattern to your MCP client](#connect-pattern-to-your-mcp-client)
if you'd rather connect a client by hand.

## What Pattern Does

Pattern looks at what you need, checks the components in your own
registered design system (code or Figma) against that need, and tells the
agent whether to:

- **Use an existing component** from your design system
- **Build a custom component**, when nothing already covers the need. The
  checklist shows which requirements the closest components miss, so the
  agent builds only the gap

A free, automatic check compares a `custom_build` verdict against
every registered candidate's real name, props, and description, and
flags any real overlap it finds, so a wrong "build it from scratch"
doesn't pass by silently.

Pattern is designed for agents to use **while they are building**. An
opt-in enforcement boundary can make the check required instead of
optional, and every decision it leads to lands in an auditable ledger
you can verify later.

It exposes fourteen tools. Five are on by default -- the ones the
register → extract → recommend → build → verify path actually needs -- and
the rest reveal themselves once you need them. See [Tool
tiers](#tool-tiers).

**Core, on by default.**

- `register_design_system` — points Pattern at your design system (a
  Figma file, a components folder, or a manifest). Required before
  `recommend_component` can judge anything.
- `recommend_component` — evaluates a UI component need against your
  registered design system and returns a structured recommendation.
- `extract_requirements` — runs just the requirement-extraction step on
  its own, so you can inspect or hand-edit the checklist before
  `recommend_component` scores against it.
- `verify_component` — run after the build: checks the built file against
  the requirement checklist item by item, with quotes the server confirms
  are really in the file, and records the outcome in the receipt.
- `record_component_decision` — records what the agent actually did so
  future recommendations in the same project can take that decision into
  account.

**[Advanced](#advanced-tools), behind `PATTERN_TOOLS=full`.**
Cost/outcome tracking, and ledger
provenance/liveness. See [Advanced tools](#advanced-tools) for the full
list.

### Tool tiers

By default Pattern's `tools/list` response advertises only the five
core tools above, so a first-time agent sees a small, obvious surface
instead of all fourteen at once. Every tool still works when called
directly, tiering only changes what gets *advertised* -- so a script or
an agent that already knows a tool's name (e.g. from this README) can
still call `read_ledger` without setting
anything. Set `PATTERN_TOOLS=full` in the server's environment to
advertise all fourteen tools immediately, e.g. for the "verify and export
old decisions" or cost-tracking workflows described below.

## How it works

![How it works](docs/images/how-it-works.png)

For each `recommend_component` call, Pattern:

1. Checks whether the need is a simple primitive that doesn't need
   scoring.
2. Requires a design system registered for the `project_id`; otherwise it
   returns an error asking you to register one.
3. Turns the request into a set of specific requirements, unless a
   checklist was already supplied (see [`checklist`](#checklist)).
4. Checks each registered candidate against the requirements using the
   props, descriptions, and Figma data captured at registration.
5. Calculates how much of the requirement is covered.
6. Decides whether to use an existing component or build a custom one. On
   a custom build, the unmet requirements show what to build.
7. Returns the result as structured JSON the calling agent can act on.

Coverage is calculated by the server from the individual requirements it
checked. It does not simply trust the percentage returned by the model.

A result can also be:

- `use_existing`
- `custom_build`
- `no_candidates_found`
- `skip_list`
- `ledger_cache_hit` — served from a recent, matching prior judgment
  instead of a fresh search+score (see
  [Per-project judgment ledger](#per-project-judgment-ledger)).

`no_candidates_found` is kept separate from a low-coverage result. Not
finding a candidate is different from finding candidates that don't cover
the requirements.

If `project_id` is supplied, Pattern also checks for past confirmed
decisions on that project and factors them in as a consistency signal —
never a rule that overrides a genuinely better match found in the current
search. Separately, `project_id` also enables the judgment ledger: a
high-confidence prior judgment matching this exact
component_need/domain/framework/existing_stack, recorded recently enough,
can be served directly (`ledger_cache_hit`) instead of running a fresh
search+score. This is the one deliberate exception to "every recommendation
searches and scores again" — see
[Per-project judgment ledger](#per-project-judgment-ledger) for the exact
rules and why it's safe.

Every result includes `computed_at`, because coverage is a snapshot of the
search at that point in time, not a permanent fact. Every result also
includes `_meta` — the timing and token cost of that specific call (see
[Cost](#cost)).

### Boundary-risk checks

The same evidence can sometimes be judged slightly differently between
model runs. When a result is close enough to a decision threshold that it
could change the verdict, Pattern automatically runs the judgment two more
times and uses the majority result.

If the three runs disagree, Pattern returns:

```json
{
  "confidence": "low",
  "ensemble": {
    "triggered": true,
    "runs": ["use_existing", "custom_build", "use_existing"],
    "agreement": "2/3"
  }
}
```

Results that are clearly inside a threshold don't trigger extra runs — see
[Cost](#cost) below for the measured impact.

### Simple primitives

These are handled locally without an API call:

| Primitive | Use it for |
| --- | --- |
| `button` | A clickable action trigger |
| `input` | A single-line text entry field |
| `checkbox` | A binary on/off toggle |
| `label` | A caption for a field or control |
| `badge` | A small status or count indicator |
| `spinner` | An indeterminate loading indicator |
| `tooltip` | A contextual hover/focus hint |
| `avatar` | A user or entity image, or initials |
| `icon` | A single glyph or symbol |

This keeps trivial requests fast and avoids unnecessary API usage.

### What powers the judgment

Each tool call makes one or more requests to the Anthropic Messages API,
using `claude-sonnet-5` by default, with a system prompt that defines the
full decision process. No web search or fetch tools are enabled: the
candidate pool is your registered design system, passed inline.

That process includes:

- Skip-list checks
- Requirement extraction
- Evidence-based coverage scoring against your registered candidates
- Decision thresholds

The model returns structured JSON. Pattern then applies important checks
itself, including recalculating coverage and applying the decision
threshold.

## Quick Start

### 1. Install and connect

```bash
npx pattern-mcp init
```

This is the only command you need to run yourself -- `npx` downloads
`pattern-mcp` on demand, then `init` detects which clients you have
installed and offers to connect each one:

- **Claude Code** -- runs `claude mcp add` for you (asks whether to make
  Pattern available in every project or just this one); skips if already
  connected (`claude mcp list` already shows it).
- **Claude Desktop** and **Cursor** -- merges a `pattern` entry into the
  client's own config file, showing the exact change before writing it
  and never touching any other server already configured there.
- **Codex CLI** -- prints the config snippet to add by hand (Codex's
  config is TOML; this doesn't auto-edit it).

Optionally pastes your `ANTHROPIC_API_KEY` into whichever configs you set
up (visible in plain text as you type it, and in the files it writes) --
press Enter to skip and add it yourself later, see
[step 2](#2-add-your-anthropic-api-key) below. Run non-interactively with
`--yes` (skips the API key prompt entirely, accepts every detected
client).

If `init` doesn't detect your client, or you'd rather set it up by
hand, see [Connect Pattern to your MCP client](#connect-pattern-to-your-mcp-client)
below for the same configs, per client, done manually. The server
command either way is `npx pattern-mcp` -- this is what your client's
config launches; you shouldn't need to run it yourself. If you do run
it bare in your own terminal (e.g. to double check the install), it
will just sit there waiting for a client and periodically remind you
to run `init` -- that's expected, not a hang.

<details>
<summary>Build from source instead</summary>

```bash
git clone <this repo>
cd pattern-mcp
npm install
npm run build
```

Use `node /absolute/path/to/pattern-mcp/dist/index.js` in place of
`npx pattern-mcp` everywhere in this README, including inside `init`'s
own generated client configs.

</details>

### 2. Add your Anthropic API key

Skipped it above, or want to change it? Pattern requires:

```
ANTHROPIC_API_KEY
```

The API account associated with this key pays for the requests Pattern
makes (see [Cost](#cost) below).

`TYPESAFE_API_KEY` (optional) turns on Jev scoring, which is faster and
cheaper than the Anthropic scorer -- see
[Scoring your design system with Jev](#scoring-your-design-system-with-jev).
With both keys set, Jev picks the match and Anthropic only runs when Jev
finds nothing that fits.

You get the key from the Anthropic Console under Settings → API Keys.
API billing is separate from Claude.ai or Claude Code subscriptions. A
Claude Pro or Max subscription does not include API usage.

### Connect Pattern to your MCP client

[Step 1](#1-install-and-connect) above (`npx pattern-mcp init`) does
this automatically for every client it detects -- the sections below
are the same configs done by hand, for a client `init` didn't detect,
or if you'd simply rather edit the config yourself.

#### Claude Code

You can add Pattern to your project's `.mcp.json` or register it with the
CLI.

For the current project:

```bash
claude mcp add pattern \
  -e ANTHROPIC_API_KEY=sk-ant-... \
  -- npx pattern-mcp
```

This uses the default local scope, so the server is available to the
current project.

To make Pattern available across your projects:

```bash
claude mcp add pattern \
  -e ANTHROPIC_API_KEY=sk-ant-... \
  --scope user \
  -- npx pattern-mcp
```

**Important:** put `-e`/`--env` and `--scope` before the `--`. Everything
after `--` is treated as the command and its arguments.

Check the connection with:

```bash
claude mcp list
```

You should see Pattern with a `✔ Connected` status.

`claude mcp add` stores the configuration in `~/.claude.json`. Avoid
`claude mcp get pattern` when possible because it can print your API key
in plaintext.

#### Cursor

Add Pattern to:

```
.cursor/mcp.json
```

#### Codex CLI

Pattern can be configured globally in:

```
~/.codex/config.toml
```

or at the project level in:

```
.codex/config.json
```

Use the MCP configuration format supported by your Codex CLI version.

#### Claude Desktop

Add Pattern through Claude Desktop's MCP settings.

The configuration looks like:

```json
{
  "mcpServers": {
    "pattern": {
      "command": "npx",
      "args": ["pattern-mcp"],
      "env": {
        "ANTHROPIC_API_KEY": "sk-ant-..."
      }
    }
  }
}
```

Restart your MCP client after adding Pattern.

Then ask your agent to list its available MCP tools and look for:

```
recommend_component
```

## Recommended workflow

The order that gets the most out of Pattern, with the call that carries
each piece of context:

1. **Register once per project** -- `register_design_system({ project_id,
   figma_file_key | figma_json_path | directory_path | manifest_path })`.
   For a Figma file, scope it with `figma_pages` / `figma_exclude_pages`;
   components-mode registration also stores the raw Figma facts (sizes,
   spacing, nested components) for later steps. Run `npx -p pattern-mcp pattern doctor`
   first if anything about setup is unclear.
2. **Extract the checklist** -- `extract_requirements({ component_need,
   domain, project_id })`. With a `project_id` that has a Figma
   registration, the checklist uses the file's real values and every item is
   tagged `figma-evidenced` / `inferred` / `general-practice`. Review or
   edit it. Skip this step to let `recommend_component` extract its own.
3. **Recommend, with the file path** -- `recommend_component({ ...,
   project_id, checklist, file_path })`. Pass `file_path` (the file you are
   about to write) **now, before writing**: the enforcement hook and the
   receipt match on it, and a receipt created after the fact is only a
   retroactive record.
4. **Build.** On `custom_build`, the `met: false` items are the gap to build;
   on `use_existing`, reuse the named design-system component.
5. **Verify** -- `verify_component({ project_id, file_path })`. Treat `fail`
   items as work left, `unverified` items as needing a human look, and
   `general-practice` items (accessibility, keyboard behavior) as things a
   design file can never prove -- expect them to stay unmet against a
   Figma-only registration.
6. **Record** -- `record_component_decision(...)` when you acted on the
   verdict (optional; it feeds the project's decision memory).

Things that will bite you if skipped: put `ANTHROPIC_API_KEY` (and
`FIGMA_ACCESS_TOKEN` / `TYPESAFE_API_KEY` if used) in the MCP server's own
`env` block, not just your shell; never paste a token into chat; restart the
client after `init` so the hook loads.

## Try it

Just ask for the UI, the way you would anyway:

> Build me a price breakdown showing nightly rate, cleaning fee, service
> fee, and taxes. I'm building an Airbnb-style booking checkout in React
> with Tailwind.

You do not need to mention Pattern. When the agent is about to build or
prototype a non-trivial component, page or screen (especially from a Figma
link, mockup or screenshot), it should check the design system first. Two
things make that happen: the server sends the agent short usage instructions
when it connects (set `PATTERN_NO_INSTRUCTIONS=1` to turn that off), and
`npx pattern-mcp init` offers to install a Claude Code skill at
`~/.claude/skills/pattern/SKILL.md`. It also offers to update an older Pattern
skill, which matters if you installed one before 0.18: that one still tells the
agent to search shadcn/ui and Mobbin. Models can still ignore instructions, so
if you want it guaranteed, turn on the [enforcement boundary](#enforcement-boundary-hook--ci-gate).

The agent should use the result to make the next decision:

- Use the recommended component from the design system, or
- Start a custom build that covers the requirements in
  `requirements_checked` marked `met: false`.

For an existing component, `component_description` explains what it does,
grounded in the data captured when you registered the design system.

## Validation examples

Pattern's validation suite uses five UI needs from an Airbnb-style rental
marketplace:

- Price breakdown with fees and taxes
- Cancellation policy display
- Host earnings dashboard
- Property image gallery
- Host-guest messaging inbox

Together, these cover different outcomes, including clear matches,
false-positive-prone searches, no candidates, and decisions close to the
threshold.

## Enforcement boundary: hook + CI gate

**Decisions can be enforced, not just tracked.** An opt-in `PreToolUse`
hook can block a new component from being written until a matching
ledger entry exists; a paired GitHub Action can also fail the PR if that
decision record isn't committed alongside the code.

**The gap this closes:** SKILL.md instructs the calling agent to call
`recommend_component` before scaffolding a new, non-trivial UI component,
but nothing before this feature *enforced* that -- an agent could simply
skip the call, and nothing server-side would know. This is opt-in and
Claude-Code-specific for the hook half; a consuming repo that never wires
either piece up gets Pattern exactly as it worked before, and any other
MCP host (Cursor, Codex, etc.) is entirely unaffected either way.

**Set it up with one command:**

```bash
npx -p pattern-mcp pattern-check-gate init
```

Confirms each step independently rather than one blanket "proceed?", and
never auto-commits -- review with `git status`/`git diff` and commit
yourself when ready:

1. Confirms a project id (pre-filled from `package.json`'s `name`, or
   your git remote/directory name -- accept it or type your own).
2. Writes or merges `.claude/settings.json` -- if one already exists, it
   parses it, leaves any unrelated hooks untouched, and only appends the
   `PreToolUse` entry if it isn't already there (safe to rerun).
3. Writes `.github/workflows/pattern-gate.yml`, if a GitHub remote is
   detected and the file doesn't already exist with different content
   (never silently overwritten).
4. Asks, as its own explicit yes/no: **mark the check required in branch
   protection?** Needs `gh` installed and authenticated with admin rights
   on the repo; skips with clear next steps otherwise. Deliberately only
   offered when no branch protection exists yet on the default branch --
   GitHub's branch-protection API replaces the *entire* configuration on
   write, not just the required-checks list, so this refuses to guess at
   merging into whatever you already have rather than risk silently
   dropping an unrelated setting (e.g. required PR reviews). If
   protection already exists, add `pattern-gate` to it by hand instead.

Run non-interactively with `--yes` (accepts every safe default; branch
protection is never auto-confirmed even then -- it's the one step that
reaches outside your local filesystem into real, shared GitHub config).

**You don't have to find this section to learn this exists.** Every
`npx pattern-mcp` run surfaces it at the same first-run moment as the
[telemetry notice](#telemetry):

- **Always**, in every context, including when a real MCP client has
  spawned this as a subprocess: a one-time, non-blocking stderr mention
  that the enforcement boundary exists and the command above sets it up.
  Same "prints once, gated by a marker file" discipline as the telemetry
  notice -- tracked at `~/.pattern/enforcement_notice_shown`
  (`PATTERN_ENFORCEMENT_NOTICE_PATH` to override), never repeats after
  that regardless of whether you act on it.
- **Only when stdin is a real terminal** (`process.stdin.isTTY`) --
  meaning a human ran `npx pattern-mcp` bare in their own shell, never
  true for a real MCP client's spawned subprocess -- it also offers a
  genuine interactive prompt right there: *"Set it up now?"* A yes runs
  the exact same `init` flow described above. The same JSON-RPC-channel
  constraint that rules out an interactive telemetry prompt (see
  [Telemetry](#telemetry)) applies here too, which is why this only ever
  asks when nothing is piping protocol messages into stdin to begin
  with.

Set `PATTERN_NO_ENFORCEMENT_NOTICE` to suppress both halves. See
`offerEnforcementSetupOnce` in `src/init-enforcement.ts` for the
implementation.

**Or set it up by hand**, two pieces, neither installed automatically:

- **`.claude/settings.json`** wired to run `npx --yes
  pattern-check-gate-hook` on `PreToolUse` (see
  `templates/claude-settings/settings.json` for the exact shape) -- a
  Claude Code hook that runs on `Write`/`Edit` calls. For a genuinely new
  `.tsx`/`.jsx` file that exports a non-trivial component, it looks up a
  ledger entry (via `~/.pattern/ledger.jsonl`, same as everywhere else in
  Pattern) whose `file_path` matches the file being written. A match
  writes a receipt and allows the write; no match blocks it with a reason
  fed back to the model as retryable guidance, not a hard failure.
  **This is the one new exception where Pattern writes into your repo**
  (`.pattern/receipts/<feature_id>.json`) -- everything else described in
  this README is read-only. `project_id` no longer needs to be set by
  hand either -- it's derived the same way `init` pre-fills it (see
  `src/project-id.ts`); set `PATTERN_PROJECT_ID` only to override that.
- **`templates/github-workflows/pattern-gate.yml`** -- a required PR
  check that reads the same receipt files back out of the diff. It never
  touches `~/.pattern/` (not reachable from a CI runner) and needs no
  `GITHUB_TOKEN` -- it trusts the committed receipt as the artifact of
  record, the same way it would trust a committed test fixture.

Receipts are `schema_version: 1` at creation. `verify_component` upgrades the
receipt for the file to `schema_version: 2` by adding a `verification` block;
the CI check matches on `file_path` alone, so v1 and v2 receipts are both
accepted.

**Components added inside existing files.** The check above looks at new
files, so a component declared inside a file that already existed used to
slip through. The CI check now also diffs each *modified* `.tsx`/`.jsx` file
against the base branch and treats every PascalCase component name that
did not exist before (a function, a `const X = () =>` / `memo` / `forwardRef`,
or a class that extends something) as needing a decision. It is name-based,
not size-based, so it has no threshold to tune. A new name is covered by
either:

- a receipt for that file whose `components` list names it. Mint one with
  `pattern-check-gate write --file <path> --components Pill,Tag` after
  `recommend_component` was called with that `file_path`; or
- an explicit skip, a comment on the line above the declaration:
  `// pattern-mcp:skip reason="CSS-only variant of Hero"`. The reason is
  required, and every skip is shown in the PR as a workflow annotation and a
  job-summary table, so an exemption is visible to a reviewer. A file-level
  `// pattern-mcp:override reason="..."` skips every new component in that file.

Limits: a component rewritten under the *same* name is not detected (the name
is not new), and detection is line-start pattern matching, not a parser.
Workflows installed before this change keep working: without `--base` the
modified files are simply not inspected. Copy the new
`templates/github-workflows/pattern-gate.yml` over your workflow to turn it on.

The join between the two depends on `file_path` being passed to
`recommend_component`/`record_component_decision` -- if it's omitted, the
gate has nothing to match against and fails closed (blocks) rather than
guessing. Pass `file_path` whenever you know it.

An escape hatch exists for both a whole-hook kill switch
(`PATTERN_NO_ENFORCEMENT_HOOK`, local only -- does not affect the CI
check) and a per-file override (a `// pattern-mcp:override reason="..."`
comment) -- the override still writes a receipt recording
`manual_override: true` and the reason, so it stays visible rather than
silent. See `src/component-gate.ts`, `src/gate-receipt.ts`,
`src/check-gate.ts` (the `pattern-check-gate` CLI, this project's first
entry point separate from the stdio MCP server), `src/check-gate-hook.ts`,
and `src/init-enforcement.ts` for the implementation, and BACKLOG.md's
"Enforcement boundary" entries for the fuller design writeup.


## Tool: `recommend_component`

### Input

```json
{
  "component_need": "price breakdown with fees and taxes",
  "domain": "Airbnb-style rental marketplace",
  "framework": "React + Tailwind",
  "existing_stack": "already using shadcn/ui",
  "project_id": "my-booking-app"
}
```

`component_need` should describe the actual UI you need, not just a
category.

Good: `price breakdown with fees and taxes`
Too vague: `pricing`

Vague requests can produce misleading matches. For example, a generic
SaaS pricing table may look like a match for "pricing" even though it
doesn't work for a booking checkout.

#### `project_id`

`project_id` is optional.

When provided, Pattern can use decisions previously recorded for the
same project (see [Per-project decision memory](#per-project-decision-memory))
as a consistency signal.

A previous decision can help the model stay consistent with similar UI
decisions, but it cannot override a better match found in the current
search.

Pattern still searches and scores every request from scratch. Past
decisions never cause a search to be skipped.

If you leave out `project_id`, Pattern does not use project memory.

#### `checklist`

`checklist` is optional -- an array of requirement strings.

When provided, `recommend_component` skips its own internal requirement
extraction entirely and scores coverage against exactly the items you
passed, instead of extracting its own checklist. Search and scoring still
run fresh every call; only the extraction step is skipped.

This is meant to be used together with [`extract_requirements`](#tool-extract_requirements):
call `extract_requirements` first, inspect (or hand-edit) the checklist it
returns, then pass that checklist here. That gives you a chance to catch a
misread requirement before Pattern spends its search+score budget.

Leave `checklist` out to keep today's default behavior: `recommend_component`
extracts its own checklist internally, exactly as before this option
existed.

#### `feature_id`

`feature_id` is optional -- a stable identifier for the feature this
component need belongs to (e.g. a ticket id or branch name). Its only use
is joining this call's cost with a later
[`report_build_cost`](#tool-report_build_cost) call for the same feature.
Omit it to have one derived deterministically from `project_id` +
`component_need`; only meaningful together with `project_id`. See
[Feature cost attribution](#feature-cost-attribution).

#### `file_path`

`file_path` is optional -- path (relative to `PROJECT_ROOT`) where this
component decision is expected to be implemented, if already known.
Usually not known yet at call time, since the decision typically precedes
the file existing. When provided, it's stored on the resulting ledger
entry and [`check_ledger_liveness`](#tool-check_ledger_liveness) can later
confirm the file still exists and still references `chosen_candidate`. It
cannot currently be attached to an entry after the fact -- see [Ledger
integrity and decision provenance](#ledger-integrity-and-decision-provenance).

**Is the checklist actually skipped, not just re-derived?** Checked, not
assumed. `breakdown_ms.extract` for a `checklist`-provided call is smaller
than the default path's, but not near-zero -- which raised the question of
whether the model is still doing some of the extraction work in that
window rather than treating the checklist as fixed input. Reading the
model's actual reasoning (via `thinking` with `display: "summarized"`,
5 runs: 3 with `checklist` provided, 2 default) answered it: the
`checklist`-provided runs' pre-search reasoning was a short, generic
"search shadcn/ui and 21st.dev" thought with no mention of the checklist's
content, e.g. *"I should look for existing image gallery component options
on shadcn/ui and 21st.dev"* -- consistently ~3-4 seconds. The default
runs' reasoning, by contrast, explicitly enumerated and derived the
checklist items (*"...mapping out the checklist: a photo grid with hero
and thumbnails... a full-screen lightbox with next/prev navigation,
keyboard support..."*) and took roughly 2x longer (~7-8 seconds). The
remaining time in the `checklist`-provided path is baseline model latency
before it decides to search, not re-extraction -- it doesn't scale with or
reference the checklist's content.

### Output

```json
{
  "verdict": "use_existing | custom_build",
  "confidence": "high | medium | low",
  "reason": "scored | no_candidates_found | skip_list",
  "computed_at": "2026-08-23",
  "requirements_checked": [
    {
      "requirement": "...",
      "met": true,
      "evidence": "..."
    }
  ],
  "coverage": "5/7 (71%)",
  "recommendation": {
    "source": "design_system | null",
    "install_command": "string | null",
    "component_description": "string | null",
    "reference": null
  },
  "ensemble": {
    "triggered": false
  },
  "checklist_source": "extracted | provided",
  "_meta": {
    "total_ms": 41516,
    "breakdown_ms": { "extract": 5006, "search": 0, "score": 33396 },
    "tokens_used": { "input": 8400, "output": 620 },
    "estimated_cost_usd": 0.14
  }
}
```

The `past_decision_signal` field is included only when there is a
relevant previous decision for the supplied `project_id`.

`checklist_source` is always present: `"extracted"` when Pattern derived
the checklist itself (the default, unchanged behavior), `"provided"` when
you passed one in via `checklist`.

`_meta` is always present. See [Cost](#cost) for what each field means,
how `breakdown_ms` is measured, and what it means when the ensemble
triggers.

### Custom builds

`recommendation.reference` is always `null`. Pattern no longer searches
Mobbin or Figma Community. On a `custom_build` verdict, read
`requirements_checked`: the items with `met: false` are what the closest
registered candidates do not cover, and the evidence text says why.

### Installation commands are not trusted

The `install_command` comes from search results. It is not verified
against a package registry, and Pattern does not execute it.

The calling agent should:

1. Show the command to the user.
2. Get confirmation.
3. Run it only after confirmation.

See [SECURITY.md](./SECURITY.md) for more details.

## Tool: `extract_requirements`

Runs only the requirement-extraction step `recommend_component` normally
does internally, and returns just the checklist -- no search, no scoring,
no verdict.

This is an opt-in, two-call pattern for agents that support tool search or
code-mode style tool use: call `extract_requirements` first, inspect (or
hand-edit) the checklist it returns, then pass that checklist to
`recommend_component`'s optional `checklist` input to score against it
directly, skipping `recommend_component`'s own internal extraction.

The single-call default -- just calling `recommend_component` with no
`checklist` -- is unchanged and is still the recommended path for most
callers. Reach for `extract_requirements` when you specifically want to
catch a misread requirement before Pattern spends its search+score budget,
not as a routine first step.

### Input

```json
{
  "component_need": "image gallery for a property listing",
  "domain": "Airbnb-style rental marketplace"
}
```

Same fields, same meaning, as `recommend_component`'s `component_need` and
`domain`. Optionally pass the same `project_id` you register and recommend
with: when that project has a Figma design system registered (components
mode), the extractor is shown the exact sizes, spacing and composition of the
most relevant components and writes the checklist with those real values.
There is no `framework` input here -- extraction is grounded in
the domain, not the framework, so `framework` doesn't affect the checklist
in `recommend_component` either.

### Output

```json
{
  "checklist": ["...", "...", "..."],
  "checklist_items": [
    { "item": "...", "basis": "figma-evidenced", "evidence": "size 320x224; vertical auto-layout, gap 16" },
    { "item": "...", "basis": "inferred" },
    { "item": "...", "basis": "general-practice" }
  ],
  "grounded_in": { "project_id": "my-booking-app", "candidates": ["Alert Dialog", "Button"] },
  "extraction_confidence": "high | medium | low",
  "_meta": {
    "total_ms": 6798,
    "breakdown_ms": { "extract": 6798, "search": 0, "score": 0 },
    "tokens_used": { "input": 275, "output": 302 },
    "estimated_cost_usd": 0.0036
  }
}
```

`checklist` is the plain list of items, a drop-in for `recommend_component`'s
`checklist` input. `checklist_items` tags each item with what it rests on:
`figma-evidenced` (it quotes a fact from the design file; the server downgrades
the tag to `inferred` if no evidence was shown or none is quoted), `inferred`
(derived from the need and domain), or `general-practice` (expected behavior a
design file cannot show -- accessibility roles, focus handling, keyboard
support). `grounded_in` is `null` when no stored Figma evidence was used.
In a measured run, `figma-evidenced` items were met 13 of 14 times and
`general-practice` items 0 of 6, which is the design file's limit, not a bad
match.

Typical latency is a few seconds -- one small API call with no tools
declared, versus `recommend_component`'s full search+score pipeline.

**`extraction_confidence` is a placeholder heuristic, not a validated
signal.** It's currently derived from how specific `component_need` is
(word count) -- the same "vague category name" problem the rest of this
README warns about elsewhere. It is not based on any measured correlation
with actual extraction quality. Treat `"low"` as a prompt to reread your
`component_need`, not as a calibrated confidence score. This is flagged
here as a known gap, to revisit once there's real usage data to base a
better signal on.

Trivial primitives (see [Simple primitives](#simple-primitives)) return an
empty `checklist` with `extraction_confidence: "high"` and no API call, the
same local skip-list short-circuit `recommend_component` uses.

## Tool: `record_component_decision`

Use this tool after the agent has actually acted on a component
decision.

For example, call it after:

- Installing an existing component
- Completing a custom build

Do not call it for every recommendation.

The tool only saves the decision. It does not run a judgment or make an
Anthropic API call.

### Input

```json
{
  "project_id": "my-booking-app",
  "component_need": "price breakdown with fees and taxes",
  "domain": "Airbnb-style rental marketplace",
  "action": "custom_built",
  "source": "custom",
  "timestamp": "2026-08-25T14:32:00.000Z",
  "time_saved_minutes": 25
}
```

- `project_id` is required and should be stable. A project directory
  path or project name works well.
- `action` must be `"installed"` or `"custom_built"`.
- `source` can be `"design_system"` or `"custom"`.
- `timestamp` is optional. If omitted, Pattern uses the current time.
- `time_saved_minutes` is optional -- the calling agent's own estimate,
  in minutes, of how much time this decision saved by having Pattern's
  verdict instead of researching candidates and judging fit from scratch.
  This is entirely self-reported. Pattern has no way to measure a
  counterfactual ("how long would this have taken without Pattern?"), so
  unlike `_meta` (Pattern's own real cost/latency for the call that
  produced the verdict), this number is never computed or verified --
  it's just recorded as-given. Omit it rather than guess a number to fill
  the field.

### Output

```json
{
  "status": "recorded",
  "project_id": "my-booking-app",
  "entry": { "..." }
}
```

## Tool: `verify_component`

Run it **after** you build. It checks the built file against the requirement
checklist that `recommend_component` recorded for that `file_path`, and
reports per item: `pass`, `fail` or `unverified`.

### Input

```json
{ "project_id": "my-booking-app", "file_path": "src/components/ConfirmDialog.tsx" }
```

`file_path` must be the same relative path you passed to
`recommend_component` -- that is how the ledger entry and its checklist are
found. Files over 80,000 characters are refused rather than truncated.

### How it decides

- Every checklist item is split into atomic clauses ("Escape cancels", "focus
  returns to the trigger", "visible focus ring"), and each clause needs its own
  **verbatim quote** from the file. The server checks the quote is really
  there (whitespace-insensitive); a `pass` with no real quote is downgraded to
  `unverified`.
- An item's status is derived by the server from its clauses: `pass` only if
  every clause passes, `fail` if any clause fails, otherwise `unverified`.
- A clause that must be **absent** ("no Tailwind classes", "LTR only") cannot
  be quoted, so the model names search terms and the server searches the file
  itself (comments ignored, whole-identifier match). It passes only if none
  occur.
- Up to five **divergences** from the Figma values are listed (size, spacing,
  radius, layout, composition) -- only for a Figma registration with stored
  evidence. The server keeps a divergence only if it can back it up: the Figma
  component and value are really in the evidence shown, the code snippet is
  really in the file, both sides are the same element, and the numbers actually
  differ. Text, labels, sample copy and colors are never divergences. Dropped
  ones are returned as `divergences_dropped` (reason + raw) but not stored in
  the receipt.

### Output and receipt

```json
{
  "ledger_entry_id": "...",
  "file_path": "src/components/ConfirmDialog.tsx",
  "file_sha256": "...",
  "summary": { "pass": 7, "fail": 0, "unverified": 2, "total": 9 },
  "items": [ { "item": "...", "status": "pass", "evidence": "...", "clauses": [ { "clause": "...", "status": "pass", "evidence": "..." } ] } ],
  "divergences": ["..."],
  "receipt": { "updated": true, "feature_id": "my-booking-app-confirm-dialog" },
  "_meta": { "estimated_cost_usd": 0.05 }
}
```

If the enforcement gate already minted a receipt for this file, the result is
written into it as `verification` and the receipt becomes **schema v2** (v1
receipts stay valid; verification never creates a receipt on its own). The
stored `file_sha256` lets anyone see when the file changed after it was
verified. `pattern-check-gate verify` in CI reports `verified`, `stale`,
`unverified_receipts` and `failed_items` -- informational, it does not block.
Each result is also appended to `~/.pattern/ledger_verifications.jsonl`
(`PATTERN_LEDGER_VERIFICATIONS_PATH`).

### Limits

It judges what the **code** says, not how it renders or behaves at runtime. A
quote proves a snippet exists, and per-clause quoting shrinks but does not
remove the chance that a clause is satisfied more loosely than it reads.
Results vary a little between runs (on a real 9-item component, 6-7 items
passed and the same one or two items flipped between `fail` and `unverified`
across runs); treat `unverified` as "look at this", not "fine". Divergences are
deliberately conservative: an empty list is common, and a divergence the model
could have reported but did not is not caught. One call costs a few cents (measured $0.045-0.056
on a ~7 KB component).

## Advanced tools

Not advertised by default -- set `PATTERN_TOOLS=full` to see these in `tools/list`, or call them directly by name at any time (see [Tool tiers](#tool-tiers)).

## Tool: `register_design_system`

Registers *this project's own* design system as the candidate pool
`recommend_component` scores against. **Required**: without a registration
for the `project_id`, `recommend_component` returns an error asking you to
register one. Works for your own component library, a Figma file, or a
design spec (shadcn/ui itself can be registered as one). Local and
per-project: registering for a `project_id` **replaces** any prior
registration for it. There's no shared/remote ledger, no
multi-user attribution, and no team auth in this scope -- those are
deliberately deferred to a team phase, only if this use case proves out.

### Input

Exactly one of `manifest_path` or `directory_path` is required, both
relative to the project root (`PATTERN_PROJECT_ROOT`, defaults to this
server's working directory) -- never an absolute path.

```json
{
  "project_id": "my-booking-app",
  "directory_path": "src/components"
}
```

- **`manifest_path`** -- a components manifest. Two recognized shapes:
  - A hand-authored JSON array of `{name, props, description,
    usage_example}` objects, optionally wrapped in `{"components": [...]}`.
  - A Storybook-exported `stories.json`/`index.json` file (an object with a
    top-level `entries` or `stories` map). Component names only in this
    case -- Storybook's basic export doesn't carry prop data, so candidates
    from this path start with an empty `props` list.
- **`directory_path`** -- a directory of real component source files,
  scanned recursively for `.jsx`/`.tsx`/`.js`/`.ts` files (excluding
  `node_modules`/`dist`/`build`/`.git` and `.test.`/`.spec.`/`.stories.`
  files). Each exported, uppercase-named function or const component found
  becomes a candidate -- including names in `export { A, B }` lists and
  generic components like `function List<T>(`; `export { X } from "./x"`
  re-exports are not candidates themselves (X is scanned in its own file)
  but are recorded in a `reexports` field on that file's candidates. Props are read in
  priority order from a `<Name>Props` interface/type, a `.propTypes` block,
  the component's own destructured parameters, or (last resort) the file's
  other `*Props` types. A `/** ... */` comment directly above a component's
  definition becomes its `description`; components without one get `null`. This is a heuristic scan, not a
  full parser -- a sparse or partial props list for some components is
  expected, not a bug, especially on plain JS with no prop typing at all.
- **`figma_json_path`** / **`figma_file_key`** -- a Figma file as the design
  system (see below).

### Output

```json
{
  "status": "registered",
  "registration": {
    "project_id": "my-booking-app",
    "source_kind": "directory_scan",
    "source_path": "src/components",
    "registered_at": "2026-09-03T18:04:11.201Z",
    "candidate_count": 29,
    "candidates": [
      { "name": "ReferralBanner", "props": ["code", "bonusAmount"], "description": null, "usage_example": null, "file_path": "rewards/ReferralBanner.jsx" }
    ]
  },
  "resolved": { "project_root": "/home/me/my-booking-app", "design_systems_path": "/home/me/.pattern/design_systems.json" }
}
```

`resolved` echoes where relative paths were resolved from (the server's
project root is often not the repo you are working in) and where the
registration is stored.

**Large registrations return a summary.** Up to 25 candidates the full
`candidates` list is returned as above. Above that, `registration` carries
`candidate_count`, `candidates_omitted`, a 10-name `candidates_preview`,
per-page counts (`candidates_by_page`, Figma sources) and any `warnings`
instead of the list, because a big Figma file would otherwise put roughly
10k tokens of candidates into the caller's context. Pass
`include_candidates: true` to always get the full list (or `false` to always
get the summary). The registration is stored in full either way.

Registering overwrites (does not merge with) any prior registration for the
same `project_id`. Once registered, `recommend_component` scores ONLY
against these candidates for calls with this `project_id` -- no separate
flag needed, it's automatic based on `project_id` alone, and no web tools
are enabled for the call. A `use_existing` verdict
scored this way always carries `"source": "design_system"` on the
resulting ledger entry, set server-side regardless of what the model wrote,
so `read_ledger` and `export_ledger_provenance` can match on it reliably.

This only writes local config to `~/.pattern/design_systems.json` (override
with `PATTERN_DESIGN_SYSTEMS_PATH`) -- it does not call the Anthropic API for a
manifest, and for a directory only to write capability summaries
(on by default with a key; see below).
Registration is a point-in-time snapshot, not a live link: re-run this
whenever the design system's own components change meaningfully.

### A safety net for a missed match

The model can occasionally say `custom_build`/`no_candidates_found`
against a registered design system even when a real match is sitting
right there in its own prompt -- a reading-comprehension miss over its own
known-complete candidate list, not evidence the list was actually empty.
When this happens, `recommend_component`'s response may carry a
`design_system_recall_check` field: a deterministic, zero-cost, local
keyword-overlap check (component name, props, description/usage_example
vs. `component_need`/`domain`) run automatically whenever reason is
`no_candidates_found` in this mode.

```json
{
  "verdict": "custom_build",
  "reason": "no_candidates_found",
  "design_system_recall_check": {
    "possible_missed_candidates": [
      { "name": "ReferralBanner", "shared_keywords": ["referral", "bonus"] }
    ],
    "note": "These registered design-system candidates share keywords with this component_need but were not selected as a match -- the verdict may have missed a real one. This is a weak, keyword-only signal, not proof of an actual match: double-check these candidates yourself (or re-run this call) before trusting custom_build here."
  }
}
```

This never overrides the verdict -- a shared keyword is weak evidence, not
proof of a real match -- it only surfaces the risk so you (or the calling
agent) know to double-check before accepting a `custom_build` verdict at
face value. Absent entirely when there's no overlap, or outside
design-system mode.

### Figma as the design system

> **Stored evidence.** In components mode, registration also keeps the raw
> facts of every component -- size, auto-layout direction/gap/padding/radius,
> the nested components it uses (named by their component set, e.g. `Button`),
> literal text, and a trimmed layer tree -- read from the component's first
> variant. They live in `figma_evidence/<project_id>.json` next to
> `design_systems.json` (`PATTERN_FIGMA_EVIDENCE_DIR` to move it), replaced on
> every re-registration and removed if you re-register from a non-Figma source.
> Frames-mode registrations have none. Scoring and extraction show the facts for
> the three best-matching candidates (`PATTERN_FIGMA_EVIDENCE_TOPK`, default 3,
> `0` turns it off); `get_figma_evidence` reads them directly. Registrations
> made before this existed need a re-register to get them.

Point registration at a Figma file instead of code: `figma_json_path` (a saved
`GET https://api.figma.com/v1/files/<file_key>` response, relative to the
project root -- fully local, no token, no network) or `figma_file_key` (fetched
from `api.figma.com` with `FIGMA_ACCESS_TOKEN` from your **environment only**,
never a tool argument, so it can't land in logs or transcripts).

- **One candidate per component set** (Figma's component-with-variants) and per
  standalone component, carrying its page/section, its Figma description, its
  variant options (e.g. `State: Default | Hover | Disabled`) and its
  boolean/text/slot property names. Hidden components (names starting with `.`
  or `_`) and instances are skipped. Two components with the same name on
  different pages stay separate.
- **Scored like any registration:** the default scorer sees the extra text, and
  with Jev scoring the same collapse-and-score path is used
  (`design_system_match.file` is `null` -- a Figma component isn't a file).
  Jev scores one entry per **page** (component family): real design systems
  define sub-parts as separate component sets (`SheetHeader`, `Table cell`), and a
  header alone can't "fully satisfy" a slide-in panel although its page is right.
  `PATTERN_JEV_FIGMA_GROUP=off` restores one entry per component set. Variant
  options and toggle names are **not** sent by default (no accuracy gain, ~50%
  more tokens); `PATTERN_JEV_FIGMA_VARIANTS=1` includes them.
- **Big files:** a real design system can be huge (the shadcn/ui Figma file is
  135 MB: fetched in ~20 s, parsed in ~3 s) and full of icon components (14,135
  of its 14,220 standalone components are icons). Use `figma_exclude_pages`
  (e.g. `["Icons"]`) or `figma_pages` to keep only what should be scored; a
  registration over `PATTERN_FIGMA_MAX_CANDIDATES` (default 3000) candidates is
  refused with the biggest pages named, not silently truncated. Registering a
  large file with vision captions can outlast a client's request timeout (some
  MCP clients cap it around 60 s); caption progress is saved as it goes, so
  re-running resumes instead of starting over.
- **No capability summaries:** there is no source file to summarize, so
  `summarize` does nothing for a Figma source.
- **Data boundary:** `figma_json_path` never leaves your machine at
  registration. `figma_file_key` sends your token and the file key to
  `api.figma.com` and downloads the whole file response. Later scoring sends
  component names, page/section names, descriptions and variant text to
  Anthropic, or to TypeSafe in Jev mode.
- **Frames mode (`figma_mode: "frames"`)** for files that never use Figma
  components -- community templates especially. One real example: the public
  "30+ Chart UI Components | BRIX Templates" file defines only 4 components (all
  style-guide helpers) and draws its 43 chart designs as plain groups named
  "Chart 1".."Chart 13" inside "Bar Charts" / "Pie Charts" frames, so the default
  mode registers nothing useful. In frames mode each named design becomes a
  candidate (a frame with 3+ substantial sub-designs is treated as a sheet, giving
  `Bar Charts > Chart 5`), and its evidence is the layer names and text inside it.
  Use `figma_pages` (e.g. `["Design"]`) to skip cover, style-guide and license
  pages. Details in the tool description.
- **What was measured (one real file, 27 needs the designs satisfy + 7 they
  don't, labels written by looking at renders of the designs -- see
  `eval/figma-eval-set.json`, `scripts/figma-frames-eval.mjs`):** components mode
  0/27; frames mode with layers + text about 18/27 right (the right design was in
  the top 4 for 26/27) but its scores overlapped the "nothing fits" scores, so
  misses came back as "not found"; names alone 1/27. Text and layer names can't
  say what is *drawn* ("candlestick", "gauge", "rings"). Adding a Claude Haiku
  **vision caption** of each rendered design as its summary took it to 23/27 right
  (27/27 in the top 4), 7/7 "nothing fits", and a clean score gap (lowest correct
  0.45-0.52 vs highest "nothing fits" 0.11-0.15), identical across three runs, for
  about $0.04 per 43 designs.
- **Vision captions (`summarize: true` with `figma_file_key`, opt-in only):**
  renders each registered design with Figma's images API and has Claude Haiku
  write a 2-sentence caption of what is drawn, stored as the candidate's
  `summary` (the default scorer and Jev both use it). **Unlike the code
  summaries this is never on by default,** because it **sends images of your
  designs to `api.anthropic.com`** and asks Figma to render them with your
  token. Needs `ANTHROPIC_API_KEY` and `FIGMA_ACCESS_TOKEN`; it is refused with
  `figma_json_path` (a saved file stays fully local and doesn't say which file to
  render). Shipped path measured live on the same file: 43 designs captioned in
  23 s for $0.039, then 23/27 right (27/27 in the top 4), 7/7 "nothing fits".
  Captions are cached by design contents: re-registering an unchanged file
  renders and sends nothing, and only a changed design is re-captioned (a
  caption survives even a re-registration without the flag). One design failing
  to caption never fails the registration. The response's `summaries.notice`
  says images were sent.
- **Components mode, measured on a real design system** (the community
  shadcn/ui Figma file: 95 candidates after excluding icons; 46 needs a
  component page satisfies, 7 nothing satisfies, 4 whose page defines no
  component at all; graded by component page; labels written by Claude from the
  page/set names, `eval/figma-shadcn-eval-set.json`, `scripts/figma-components-eval.mjs`):
  one entry per component set got ~27/46 right (the right family in the top 4 for
  ~44) with Jev under-confident, versus 40/46 for a one-call Sonnet baseline on the
  same evidence; **one entry per page got 41-42/46** (top 4: 46/46), 7/7 "nothing
  fits", at 4.3k tokens per need; the 4 pages with no component were correctly not
  found in 3 of 4. Variant text made no difference (41 vs 41; 27 vs 28) and
  vision captions added almost nothing (28 vs 27) on this well-named library, so
  **captions pay off for frames-mode template files, not for a well-named component
  library.** Score gap is thin here: lowest correct 0.40 vs highest "nothing fits"
  0.29-0.30.
- **Third file, a small atomic UI kit** ("Design System | UI kit | +6000
  Components" community file: 24 component sets, ~6,000 icons excluded, 39
  page-family candidates; 22 positive needs + 7 nothing-fits,
  `eval/figma-uikit-eval-set.json`): the shipped default got **22/22 right and
  7/7 "nothing fits" in 3 of 3 runs** at ~1.5k tokens/need, lowest correct
  0.44-0.48 vs highest "nothing fits" 0.21-0.23 (a wider gap than shadcn/ui); a
  one-call Sonnet baseline also got 22/22. This is a **ceiling result**: the pages
  are named after the components, so it shows the path holds up on a third file
  shape, not that it discriminates hard cases. Vision captions were not run here.
- **Status -- read this:** each mode has run on one to three real Figma files
  (a chart template; the shadcn/ui design system; a small atomic UI kit), all
  with labels written by Claude, not independently. Treat results on other files
  as unvalidated.

### Capability summaries (`summarize`, on by default)

Registering with `directory_path` also writes a short (2-3 sentence) capability
summary for each scanned file with Claude Haiku and stores it on the
registration -- **by default, whenever `ANTHROPIC_API_KEY` is set.** In the
eval this was the single biggest gain for the Jev scorer below (29/29 in 5 of
5 runs without prop names, lowest correct score 0.55 vs a highest "nothing
fits" of 0.18); on real libraries it scored 37/38 with the lowest correct 0.56
and highest "nothing fits" 0.12. The default scorer also sees the summaries.

- **This sends code.** Up to 8000 characters of each file that needs a summary
  go to `api.anthropic.com`. The response's `summaries.notice` says so every
  time it happens. To keep registration fully local, pass `"summarize": false`
  on a call or set `PATTERN_NO_SUMMARIES=1` to turn it off everywhere. No key,
  nothing is sent (the response says the summaries were skipped).
- **Cost:** about 0.2 cents per file (measured: 96 files, $0.18 total). Capped at
  `PATTERN_SUMMARY_MAX_FILES` files per registration (default 200); the
  response's `summaries` block reports files generated / reused / failed /
  skipped, tokens and estimated cost.
- **Cached by file content:** re-registering pays only for files that changed.
  Unchanged files keep their summary even with `summarize: false`; a changed
  file's stale summary is dropped and rewritten on the next default run.
- A file that fails to summarize is left without one -- the registration still
  succeeds. `"summarize": true` (explicit) is refused, leaving your previous
  registration untouched, when there is no key or you passed `manifest_path`
  (no source files to read); the unset default never refuses, it just skips.

### Scoring your design system with Jev

Set `TYPESAFE_API_KEY` and `recommend_component` scores your registered design
system with [Jev](https://typesafe.ai) (TypeSafe AI): sub-second, no Anthropic
call for a match. `PATTERN_SCORER` picks the mode:

| `PATTERN_SCORER` | Behavior |
|---|---|
| unset (with `TYPESAFE_API_KEY`) | **Hybrid.** Jev scores every call. A match returns immediately. If Jev finds nothing that fits, Anthropic runs (when `ANTHROPIC_API_KEY` is set) to produce the requirement-by-requirement gap list, and the result carries `jev_screen`. Without an Anthropic key you get Jev's plain "nothing fits" result. |
| `jev` | Jev only. No Anthropic call ever, so a `custom_build` has no gap list. |
| `anthropic` | Anthropic only, even when `TYPESAFE_API_KEY` is set. |

Without `TYPESAFE_API_KEY` (and `PATTERN_SCORER` unset) Pattern uses the
Anthropic scorer for everything. `extract_requirements`, calls that pass a
`checklist`, registration summaries and Figma vision captions still need
`ANTHROPIC_API_KEY`.

- **Nothing fits -> it says so.** If no component scores at least
  `PATTERN_JEV_FOUND_THRESHOLD` (default `0.4`), Jev's result is
  `verdict: "custom_build"`, `reason: "no_candidates_found"`, with a plain
  `not_found_message`. In hybrid mode with an Anthropic key, that result is
  replaced by the Anthropic pass (checklist, coverage, unmet items) and
  annotated with `jev_screen`.
- **Found -> `use_existing`**, `recommendation.source: "design_system"`, plus
  `design_system_match` (best file, runners-up, raw score). Confidence is
  `low` or `medium`, never `high`: Jev's scores rank well but are **not
  calibrated probabilities**, so treat the number as a ranking signal.
- **What is scored:** one entry per source file (exports, re-exports, the
  doc comment above each component, and a `summary` if the registration has
  one). Prop lists are intentionally not sent -- in the eval they misled the
  scorer. Large libraries are split into batches
  (`PATTERN_JEV_BATCH_TOKENS`, default 18000).
- **Not used when** you pass your own `checklist` (Jev scores whole files, not
  checklist items) -- that call takes the normal Anthropic path.
- **Cost:** reported as `0` with a `cost_note` unless you set
  `PATTERN_JEV_USD_PER_MTOK_IN` / `PATTERN_JEV_USD_PER_MTOK_OUT`.
- **Data boundary:** your `component_need` text and the evidence above go to
  `api.typesafe.ai`. Your source files are read locally and never sent whole.
  See [SECURITY.md](./SECURITY.md).

The threshold and the collapse/no-props choices came from
`scripts/design-system-jev-eval.mjs` on 38 needs over two libraries, so they
are a first guess -- re-run the eval on your own design system before
trusting them.

## Tool: `get_figma_evidence`

Returns the raw Figma facts stored at registration for a component, with no
API call.

```json
{ "project_id": "my-booking-app", "name": "Alert Dialog" }
```

Pass `node_id` (exact) or `name` (case-insensitive, exact match first, then
substring, up to 5 results). Output is `{ captured_at, matches: [{ node_id,
name, source_node, size, layout, instances, texts, tree, truncated }] }`.
Returns an error with a hint if the project has no stored evidence (non-Figma
or frames-mode registration, or one made before evidence capture).

## Tool: `read_ledger`

Lists past `recommend_component` judgments for a `project_id` -- every
call that reached the API and produced a verdict, not just ones explicitly
confirmed via `record_component_decision`. Useful for auditing what
Pattern has already judged for a project, or for understanding why a call
came back with `served_from_ledger: true`.

### Input

```json
{
  "project_id": "my-booking-app",
  "component_need": "cancellation",
  "limit": 10
}
```

- `project_id` is required.
- `component_need` is optional -- a simple keyword filter (substring
  match, no embeddings) against stored entries' `component_need`. Omit to
  list everything for the project.
- `limit` is optional, defaults to 20. Most recent entries first.
- `feature_id` is optional. When provided, `component_need` and `limit`
  are ignored and the response is a full cost rollup for that one feature
  instead of a keyword listing -- see [Feature cost
  attribution](#feature-cost-attribution).

### Output

```json
{
  "project_id": "my-booking-app",
  "entries": [
    {
      "id": "a1b2c3d4-...",
      "timestamp": "2026-08-29T19:50:47.073Z",
      "project_id": "my-booking-app",
      "feature_id": "3f9a21c0",
      "component_need": "cancellation policy display with refund tiers by date",
      "domain": "Airbnb-style rental marketplace",
      "framework": "React + Tailwind",
      "checklist": ["...", "..."],
      "checklist_source": "extracted",
      "candidates_evaluated": [
        { "source": "ReUI (reui.io)", "name": "Timeline", "url": "https://reui.io/components/timeline", "coverage_pct": 62.5 }
      ],
      "verdict": "use_existing",
      "chosen_candidate": "Timeline",
      "confidence": "low",
      "reason": "scored",
      "coverage": "5/8 (62.5%)",
      "cost_usd": 0.087,
      "cache_hit": false,
      "project_conventions_snapshot": "9f3a1c7e2b0d4f5a",
      "file_path": null,
      "snapshot_ref": "a1b2c3d4e5f6...",
      "last_verified_live": null,
      "live_status": "unknown",
      "reconstructed_snapshot_ref": null
    }
  ]
}
```

`file_path`/`snapshot_ref`/`last_verified_live`/`live_status`/
`reconstructed_snapshot_ref` are the ledger integrity + decision
provenance fields -- see [Ledger integrity and
decision provenance](#ledger-integrity-and-decision-provenance) and [Tool:
`check_ledger_liveness`](#tool-check_ledger_liveness). Entries written
before this feature shipped read back with `file_path`/`snapshot_ref`/
`last_verified_live` as `null` and `live_status` as `"unknown"` rather
than missing keys.

Passing `feature_id` instead returns:

```json
{
  "project_id": "my-booking-app",
  "feature_id": "3f9a21c0",
  "verdict_entries": [ "...same shape as above, filtered to this feature_id..." ],
  "build_records": [
    { "id": "...", "timestamp": "...", "project_id": "my-booking-app", "feature_id": "3f9a21c0", "tokens_used": 9000, "cost_usd": 1.25, "outcome": "shipped" }
  ],
  "total_cost_usd": 1.34,
  "outcome_proxy": { "time_to_merge_hours": 3.5, "reworked": true, "days_to_rework": 12, "status_at_30d": "kept" },
  "outcome_proxy_history": [ "...every raw report_outcome_proxy record for this feature_id, oldest first..." ]
}
```

`outcome_proxy` is `null` (and `outcome_proxy_history` an empty array)
when no `report_outcome_proxy` calls have been made for this feature yet
-- see [Outcome proxies](#outcome-proxies).

Each entry holds only distilled fields -- `candidates_evaluated` never
contains raw HTML, full prop tables, or the per-requirement evidence text
`recommend_component` itself returns. See
[Data minimization](#data-minimization) below.

## Tool: `report_build_cost`

Self-reports the end-to-end build cost for one feature. Pattern only ever
sees the cost of judging *what* to use (`recommend_component`'s own
`_meta.estimated_cost_usd`); everything past that -- the actual scaffold,
install, or custom build -- happens outside Pattern entirely and Pattern
has no way to observe it. Call this once, after the calling agent's build
for a feature is actually complete (shipped, abandoned, or replaced), not
on every verdict.

### Input

```json
{
  "feature_id": "3f9a21c0",
  "project_id": "my-booking-app",
  "tokens_used": 9000,
  "cost_usd": 1.25,
  "outcome": "shipped"
}
```

- `feature_id` is required -- either a value you explicitly passed to an
  earlier `recommend_component` call for this feature, or (if you didn't)
  the same value `recommend_component` derives on its own:
  `sha256(project_id + "::" + component_need, lowercased/trimmed)`
  truncated to 8 hex characters. When in doubt, call `read_ledger` with
  just `project_id` and copy the `feature_id` off the relevant entry
  rather than re-deriving it by hand.
- `project_id` is optional but recommended -- without it, this record
  still joins to a `recommend_component` entry by `feature_id` alone, but
  `read_ledger`'s rollup can't scope it to one project.
- `tokens_used` is optional.
- `cost_usd` is required -- your own real number, not Pattern's.
- `outcome` is required: `"shipped"`, `"abandoned"`, or
  `"replaced_with_existing"`.

### Output

```json
{
  "status": "recorded",
  "record": {
    "id": "c5706b47-...",
    "timestamp": "2026-09-02T01:25:29.653Z",
    "project_id": "my-booking-app",
    "feature_id": "3f9a21c0",
    "tokens_used": 9000,
    "cost_usd": 1.25,
    "outcome": "shipped"
  }
}
```

This only appends a local record to `~/.pattern/build_ledger.jsonl`
(override with `PATTERN_BUILD_LEDGER_PATH`) -- it never re-runs any
judgment and never calls the Anthropic API.

## Tool: `report_outcome_proxy`

Self-reports a value signal for one feature, deliberately independent of
Pattern's own verdict -- the whole point is a signal that could
*contradict* the verdict, so nothing on this path ever reads
`coverage_pct`, `confidence`, or any other Pattern-produced field. Compute
`reworked`/`days_to_rework` and `time_to_merge_hours` from your own repo's
real git history (e.g. `git log --follow` against the files this
feature's build touched) rather than relying on Pattern -- rework rate and
time-to-merge need real git *history*, a materially bigger surface than
the one narrow, read-only exception described in [Ledger integrity and
decision provenance](#ledger-integrity-and-decision-provenance) below.
Report `status_at_30d` only once a real ~30-day-post-merge horizon has
actually passed.

Safe to call more than once for the same `feature_id` as more signal
becomes available over time -- e.g. `time_to_merge_hours` right after
merge, `reworked` on a later re-check, `status_at_30d` at the 30-day mark.
`read_ledger`'s `feature_id` rollup merges every report into one
latest-value-per-field view (a later report only overwrites the specific
fields it includes, never the others).

### Input

```json
{
  "feature_id": "3f9a21c0",
  "project_id": "my-booking-app",
  "reworked": true,
  "days_to_rework": 12
}
```

- `feature_id` is required.
- `project_id` is optional but recommended, same reasoning as
  `report_build_cost`.
- `reworked`, `days_to_rework`, `time_to_merge_hours`, `status_at_30d` are
  all individually optional, but **at least one is required** -- an empty
  report is rejected rather than silently recording nothing.

### Output

```json
{
  "status": "recorded",
  "record": {
    "id": "8a2f1e0c-...",
    "timestamp": "2026-09-16T18:04:12.881Z",
    "project_id": "my-booking-app",
    "feature_id": "3f9a21c0",
    "reworked": true,
    "days_to_rework": 12
  }
}
```

This only appends a local record to `~/.pattern/outcome_proxies.jsonl`
(override with `PATTERN_OUTCOME_PROXY_PATH`) -- it never calls the
Anthropic API.

## Tool: `check_ledger_liveness`

Checks whether ledger entries for a `project_id` are still **live** --
does the `file_path` recorded on the entry (if any, see
[`file_path`](#tool-recommend_component)) still exist, and does it still
mention `chosen_candidate`. See [Ledger integrity and decision
provenance](#ledger-integrity-and-decision-provenance) for the full design
and its deliberate limits.

This is the **one exception** to Pattern otherwise having no filesystem
access to your repo (see [Outcome proxies](#outcome-proxies) above) --
scoped narrowly to read-only `fs.existsSync`/file-read calls against
`PROJECT_ROOT` (defaults to this server's own working directory; override
with `PATTERN_PROJECT_ROOT`). It never writes to your repo and never runs
an arbitrary shell command.

### Input

```json
{
  "project_id": "my-booking-app",
  "ledger_entry_id": "a1b2c3d4-..."
}
```

- `project_id` is required.
- `ledger_entry_id` is optional -- check just that one entry instead of
  every entry for `project_id` that has a `file_path` set.

### Output

```json
{
  "project_id": "my-booking-app",
  "checked": 1,
  "total_entries": 2,
  "results": [
    {
      "ledger_entry_id": "a1b2c3d4-...",
      "component_need": "cancellation policy display with refund tiers by date",
      "file_path": "src/components/CancellationPolicy.tsx",
      "live_status": "live",
      "checked_at": "2026-09-02T20:11:03.442Z",
      "note": null
    },
    {
      "ledger_entry_id": "e5f6a7b8-...",
      "component_need": "gallery",
      "file_path": null,
      "live_status": "unknown",
      "checked_at": null,
      "note": "no file_path recorded on this entry -- nothing to check"
    }
  ]
}
```

`live_status` is one of `"live"`, `"orphaned"`, `"unknown"`, or
`"dangling"` (only ever produced by
[`sweep_ledger_liveness`](#tool-sweep_ledger_liveness)'s cluster
detection, never by a single-entry `check_ledger_liveness` call -- see
[Ledger integrity and decision
provenance](#ledger-integrity-and-decision-provenance)).
Entries with no `file_path` are listed but never checked or written to
`ledger_liveness.jsonl` -- their status is permanently `"unknown"` since
there's nothing to check. Results here are also layered onto
`read_ledger`'s `live_status`/`last_verified_live` fields for the same
entries afterward -- `check_ledger_liveness` is the only thing that
advances those fields past their write-time defaults.

## Tool: `sweep_ledger_liveness`

Batch version of [`check_ledger_liveness`](#tool-check_ledger_liveness):
updates `live_status` for every `file_path`-bearing entry across an
entire project, or -- when `project_id` is omitted -- every `project_id`
present in the ledger. This is the "on a schedule (project open or cron)"
half of the referential-integrity design that `check_ledger_liveness`'s
on-demand, single-project call doesn't cover.

**Pattern has no daemon or scheduler of its own.** Each server invocation
is transient, tied to its MCP host's lifecycle -- there is nowhere inside
this server for a cron job to live. This tool is meant to be invoked by
whatever external scheduler you already have (a cron job, a CI step
running nightly), not something Pattern triggers automatically or ever
will on its own.

### Input

```json
{
  "project_id": "my-booking-app"
}
```

`project_id` is optional -- omit it to sweep every `project_id` present
in the ledger in one call.

### Output

```json
{
  "projects_swept": 2,
  "total_entries_checked": 14,
  "dangling_clusters": [
    { "project_id": "my-booking-app", "feature_id": "3f9a21c0", "entry_ids": ["...", "..."] }
  ],
  "per_project": [
    { "project_id": "my-booking-app", "checked": 9, "total_entries": 12, "dangling_clusters": 1 },
    { "project_id": "other-project", "checked": 5, "total_entries": 5, "dangling_clusters": 0 }
  ]
}
```

### Dangling clusters, and how "cluster" maps onto what the ledger actually stores

The ledger has no explicit entry-to-entry reference field -- each line is
an independent judgment record. `feature_id` (see [Feature cost
attribution](#feature-cost-attribution)) is the one real grouping
construct that already exists, so a "cluster" here means every entry
sharing one `feature_id`, and "no live anchor" means none of them
resolved to `live_status: "live"`. A single-entry group is just an
ordinary orphaned/unknown entry, not a cluster phenomenon, so groups of
one are never flagged.

Every entry in a qualifying cluster gets `live_status: "dangling"` --
overriding whatever `"orphaned"`/`"unknown"` value it had -- visible on
its next `read_ledger`/`check_ledger_liveness` read via the same
`ledger_liveness.jsonl` overlay `check_ledger_liveness` already writes to
(see [Referential integrity](#referential-integrity-file_path--live_status)).
Tested against the exact repro shape reported by a user: 13 entries, 12
sharing a `feature_id` with no live anchor among them, 1 separate and
live -- all 12 flag `dangling`, the 13th doesn't. Also tested at 200 and
1,000 synthetic entries without reintroducing search+score-class latency
(both complete in well under a second -- this is `fs.existsSync` calls
and in-memory grouping, not API calls).

## Tool: `export_ledger_provenance`

Formats one ledger entry -- requirements checklist, candidates compared,
verdict, confidence, `snapshot_ref` -- as a single markdown block: a
stable, portable record of that decision you can paste into a PR
description or issue by hand. See [Ledger integrity and decision
provenance](#ledger-integrity-and-decision-provenance) for the full
design and its deliberate limits.

Pure and deterministic: the same entry always produces byte-identical
markdown, since the function reads nothing but its input (no live system
time, no disk state). This only formats and returns text -- it does not
post anything anywhere; see
[`post_ledger_provenance_to_github`](#tool-post_ledger_provenance_to_github)
below for that.

### Input

```json
{
  "project_id": "my-booking-app",
  "ledger_entry_id": "a1b2c3d4-..."
}
```

Both fields are required -- unlike `check_ledger_liveness`, there's no
"every entry for this project" mode, since a provenance artifact is
inherently about one specific decision.

### Output

```json
{
  "ledger_entry_id": "a1b2c3d4-...",
  "markdown": "## Pattern decision: cancellation policy display with refund tiers by date\n\n- **Verdict:** use_existing (confidence: high)\n- **Reason:** scored\n- **Coverage:** 5/8 (62.5%)\n- **Domain:** Airbnb-style rental marketplace\n- **Framework:** React + Tailwind\n- **Snapshot:** `9f3a1c7e2b0d4f5a6b7c8d9e0f1a2b3c4d5e6f70`\n- **Judged at:** 2026-08-29T19:50:47.073Z\n\n### Requirements checked\n- ...\n\n### Candidates compared\n| Source | Name | Coverage | Chosen |\n| --- | --- | --- | --- |\n| ReUI (reui.io) | Timeline | 62.5 | ✓ |\n\n_Generated by Pattern (`export_ledger_provenance`) from ledger entry `a1b2c3d4-...`._"
}
```

Errors (as `isError: true`, not a thrown exception) when `ledger_entry_id`
doesn't match any entry for that `project_id` -- including when the id is
real but belongs to a different project, since entries are always scoped
per `project_id`.

For a `custom_build` verdict, the candidates section explains that gap in
prose instead of an empty table.
A `null` `snapshot_ref` (project root wasn't a git repository at judgment
time) renders as prose too, not the literal word `null`.

## Tool: `backfill_ledger_snapshot_ref`

Best-effort reconstruction of `snapshot_ref` for ledger entries written
before that field existed (or written outside a git repository): finds
the commit that was `HEAD` at or just before each entry's own timestamp
(`git log --before=<timestamp> -1 --format=%H`). Entries that already
have a real `snapshot_ref` are reported but never touched -- backfill
only ever fills a gap, never second-guesses a captured value.

### Input

```json
{
  "project_id": "my-booking-app",
  "ledger_entry_id": "a1b2c3d4-..."
}
```

`ledger_entry_id` is optional -- omit it to backfill every entry in the
project missing `snapshot_ref`.

### Output

```json
{
  "project_id": "my-booking-app",
  "attempted": 3,
  "reconstructed": 2,
  "results": [
    { "ledger_entry_id": "a1b2c3d4-...", "already_had_snapshot_ref": false, "reconstructed_snapshot_ref": "9f3a1c7e2b0d4f5a6b7c8d9e0f1a2b3c4d5e6f70" },
    { "ledger_entry_id": "e5f6a7b8-...", "already_had_snapshot_ref": false, "reconstructed_snapshot_ref": null }
  ]
}
```

### A reconstructed value is always labeled, never presented as real

Necessarily an approximation, not a guarantee: a rebase, force-push, or
history rewrite since that timestamp can make "the commit `HEAD` pointed
to then" no longer resolve to what the codebase actually looked like at
judgment time. Every attempt is persisted (including failures -- a
project whose git history doesn't reach back that far, or that isn't a
git repository at all) to `~/.pattern/snapshot_backfill.jsonl` (override
with `PATTERN_SNAPSHOT_BACKFILL_PATH`), and surfaces on later reads as
`reconstructed_snapshot_ref` -- a field kept fully separate from
`snapshot_ref` itself, never overwriting or being confused with it.
[`export_ledger_provenance`](#tool-export_ledger_provenance) and
[`post_ledger_provenance_to_github`](#tool-post_ledger_provenance_to_github)
both render a reconstructed value with an explicit "(reconstructed via
backfill -- best-effort approximation, not the original captured
snapshot)" label, never silently as if it were equivalent to a value
captured live.

Tested against a real throwaway git repo with known commit history (an
entry timestamped between two real commits reconstructs to exactly the
first one), a 200-entry synthetic ledger outside any git repo (every
attempt fails fast and reports `null` rather than throwing), and a
read-only run against this project's own real `coop-commerce` ledger
entries, per the spec's own test plan.

## Tool: `post_ledger_provenance_to_github`

Posts one ledger entry's provenance artifact (the same content
`export_ledger_provenance` produces) as a real comment on a GitHub PR or
issue. **This is the one tool in this server with a real, visible side
effect on a third-party service** -- every other tool here only ever
touches local files. Confirm with the user before calling it, the same
way you're expected to confirm before running a suggested
`install_command` (see [Installation commands are not
trusted](#installation-commands-are-not-trusted) and SECURITY.md).

GitHub treats a PR and an issue identically for comments (both use the
same `/issues/{number}/comments` endpoint), so there's one input shape
for both -- no separate "is this a PR" flag.

### Auth: `GITHUB_TOKEN`, not a GitHub App

This resolves the open question left in [Ledger integrity and decision
provenance](#ledger-integrity-and-decision-provenance)'s earlier writeup
in favor of a **personal access token**, read from the `GITHUB_TOKEN`
environment variable -- the same convention every GitHub Action and the
`gh` CLI itself already use. Needs `repo` scope. A GitHub App was the
alternative on the table, but it needs a hosted installation flow and a
webhook receiver, which contradicts this project's entire distribution
model (a local npm package, no hosted infrastructure -- see [Ledger
integrity and decision provenance](#ledger-integrity-and-decision-provenance)
and the Pattern Primer's build-order principle). Pattern manages no
GitHub credential of its own, the same way it manages no git credential
for `snapshot_ref` -- it just reads what's already in your environment.

### Input

```json
{
  "project_id": "my-booking-app",
  "ledger_entry_id": "a1b2c3d4-...",
  "repo": "my-org/my-booking-app",
  "issue_number": 42
}
```

All four fields are required.

### Output

```json
{
  "posted": true,
  "comment_url": "https://github.com/my-org/my-booking-app/pull/42#issuecomment-...",
  "comment_id": 123456789
}
```

### Idempotent by construction

Every posted comment is prefixed with a hidden HTML marker keyed to the
ledger entry's id (`<!-- pattern-ledger-provenance:<id> -->`). A call
first checks the thread's existing comments (most recent 100 -- full
pagination isn't handled yet) for that marker; if found, it returns
`{ "posted": false, "reason": "already_posted", "comment_url": "..." }`
pointing at the existing comment instead of creating a duplicate. A
repeat call is always safe to make.

Errors (`isError: true`) clearly on: no `GITHUB_TOKEN` set, a malformed
`repo` (not `owner/repo`), an unknown `ledger_entry_id`, or a GitHub API
error (bad credentials, repo/issue not found, rate limit) -- the error
message includes the real HTTP status and GitHub's own error text.

## Feature cost attribution

Every `recommend_component` call that writes to the ledger -- a fresh
judgment *or* a $0 [ledger cache hit](#the-cache-hit-exception) -- now
carries a `feature_id`, plus its own `cost_usd` and `cache_hit`. Pair that
with `report_build_cost`'s build-time record and `read_ledger`'s
`feature_id` rollup, and total spend on a feature (judgment + build,
across however many calls) is queryable end to end, not just the cost of
one verdict call.

`feature_id` defaults to a deterministic derivation --
`sha256(project_id + "::" + component_need)` truncated to 8 hex chars --
so repeat calls for the same feature land under the same id automatically,
with no coordination needed between `recommend_component` and
`report_build_cost` calls. Pass your own `feature_id` explicitly (e.g. a
ticket id or branch name) if you'd rather key on something stable on your
own side.

## Outcome proxies

Cost data alone (`feature cost attribution` above) can't answer whether a
cheaper build was actually *worth it* -- comparing it against Pattern's
own verdict/`coverage_pct` would be circular, since that's the very thing
being evaluated. `report_outcome_proxy` attaches a cheap, non-circular
value signal per `feature_id` instead:

- **`reworked` / `days_to_rework`** (primary proxy) -- was any file this
  feature's build touched modified again after the original merge, and if
  so, how soon? Computed from real git history, not Pattern's own data.
- **`time_to_merge_hours`** (secondary proxy) -- how long the feature
  took from first commit to merge.
- **`status_at_30d`** (tertiary, longer-horizon proxy) -- at a ~30-day
  horizon, does the component Pattern recommended still exist in the
  codebase, unchanged in kind (`"kept"`), was it swapped for a different
  approach (`"replaced"`), or removed entirely (`"removed"`)?

`read_ledger`'s `feature_id` rollup returns both `outcome_proxy` (the
merged latest-value-per-field view) and `outcome_proxy_history` (every
raw report, in case the timeline itself matters) alongside the cost
figures from [Feature cost attribution](#feature-cost-attribution) above
-- so "what did this feature cost end to end, and did it hold up?" is
answerable from one `read_ledger` call.

## Per-project judgment ledger

Distinct from [per-project decision memory](#per-project-decision-memory)
below -- that file only gains an entry when `record_component_decision` is
explicitly called. The ledger instead gains one entry automatically for
**every** `recommend_component` call with a `project_id` that lands on
reason `"scored"` or `"no_candidates_found"` -- whether that's a fresh
call that reached the API, or a $0 [ledger cache
hit](#the-cache-hit-exception) served without one (`cache_hit: true`,
`cost_usd: 0`), so a feature's total cost still rolls up correctly even
once most of its later calls are free. See [Feature cost
attribution](#feature-cost-attribution).

Pattern stores it locally in:

```
~/.pattern/ledger.jsonl
```

Change the location with `PATTERN_LEDGER_PATH`. One JSON object per line
(append-only, JSONL).

### The cache-hit exception

Every other part of Pattern scores fresh every time (see
[No caching, by design](#no-caching-by-design)). The ledger is the one
deliberate exception: a later `recommend_component` call with a matching
`project_id` **can** be served directly from a prior entry, skipping
search+score entirely, when **all** of the following hold:

- `component_need` matches exactly (case-insensitive).
- `domain` and `framework` match exactly.
- `existing_stack` hashes to the same value as the stored entry's
  (both omitted counts as a match).
- The stored entry's `confidence` is `"high"`.
- The stored entry's `reason` is `"scored"` or `"no_candidates_found"`.
- The stored entry is no older than `PATTERN_LEDGER_TTL_DAYS` (default
  **30** days, configurable).

When served this way, the response has `reason: "ledger_cache_hit"`,
`served_from_ledger: true`, `ledger_entry_id`, and
`original_verdict_timestamp` -- so nothing is ever silently passed off as
freshly verified. `_meta.estimated_cost_usd` and `tokens_used` are
genuinely `0`: no API call happened. `requirements_checked` is `null` on
this path -- the ledger never stores per-requirement evidence text (see
[Data minimization](#data-minimization)), so a cache hit can only replay
the verdict/confidence/coverage/chosen-candidate, not the original
per-requirement reasoning.

Any mismatch on the criteria above -- a different `domain`, a changed
`existing_stack`, an entry that's gone stale, or one that wasn't
high-confidence -- falls through to a normal, fresh search+score call.

### Turning the cache-hit exception off

Set `PATTERN_NO_LEDGER_CACHE_HIT` (any truthy value) to restore
"every `recommend_component` call always scores fresh" without removing
any ledger code. This disables only the cache-hit short-circuit --
entries are still written to `ledger.jsonl` and `read_ledger` still works
either way, so the audit trail keeps growing even with the switch on.
Unset the variable to re-enable cache hits again at any time.

### Data minimization

Nothing written to the ledger ever contains raw search/fetch content.
Every candidate is reduced to exactly four fields before it's written --
`source`, `name`, `url`, `coverage_pct` -- enforced at the type level
(`assertDistilledCandidateShape` in `src/index.ts`), not just by
convention: a raw or extended object throws rather than silently
persisting. Run `node scripts/verify-ledger-boundary.mjs` (after
`npm run build`) to check this boundary directly.

## Ledger integrity and decision provenance

Two gaps in the ledger, surfaced from user feedback: it tracks that a
decision was made, but not whether the thing it decided about is still
live in your codebase, and it stores the checklist/verdict but not a
version pin or an exportable artifact you can attach to a PR or issue.
Both are now fully addressed, across five tools -- **old decisions can be
checked, not just logged**: [`check_ledger_liveness`](#tool-check_ledger_liveness)
verifies that the file where a decision was implemented still exists and
still uses the recommended component, marking it an orphaned entry if it
doesn't ([`sweep_ledger_liveness`](#tool-sweep_ledger_liveness) is the
batch/scheduled version of the same check); **decisions can become
shareable records**: [`export_ledger_provenance`](#tool-export_ledger_provenance)
turns a decision into a self-contained Markdown record, and
[`post_ledger_provenance_to_github`](#tool-post_ledger_provenance_to_github)
can attach it directly to the relevant PR or issue; and **older decisions
aren't left behind**: [`backfill_ledger_snapshot_ref`](#tool-backfill_ledger_snapshot_ref)
adds a `snapshot_ref` to decisions created before this feature existed, so
the liveness check above works retroactively. See
`pattern-ledger-integrity-and-provenance-spec.md` for the original phased
plan this was built against.

**This required the one deliberate exception** to Pattern otherwise having
[no filesystem/git access to your repo](#per-project-judgment-ledger) at
all (the principle `report_build_cost`/`report_outcome_proxy` are built
around). Still narrow, still all read-only, and still nothing here ever
writes to your repo or runs an arbitrary git/shell command:

- `git rev-parse HEAD`, on every ledger write, to capture `snapshot_ref`.
- `git log --before=<timestamp> -1 --format=%H`, only inside
  `backfill_ledger_snapshot_ref`, to reconstruct a best-effort
  `snapshot_ref` for an entry that predates it.
- `fs.existsSync` plus a plain-text read of one file, only for a
  `file_path` you explicitly passed to `recommend_component`, only inside
  `PROJECT_ROOT` (see below) -- what
  [`check_ledger_liveness`](#tool-check_ledger_liveness)/[`sweep_ledger_liveness`](#tool-sweep_ledger_liveness)
  check. `post_ledger_provenance_to_github` additionally makes a real,
  visible network call to the GitHub API -- see that tool's own docs,
  it's a materially different kind of exception (a third-party service,
  not your local machine) from the four above.

### `PROJECT_ROOT`

Defaults to `process.cwd()` -- for a locally-run stdio MCP server, that's
normally the consuming repo's root, since MCP hosts typically launch the
server with the project directory as its working directory. Override with
`PATTERN_PROJECT_ROOT` if that assumption doesn't hold for your setup.

A `file_path` that's absolute or escapes `PROJECT_ROOT` via `../` resolves
to `live_status: "unknown"` rather than being read -- belt-and-suspenders,
since the calling agent already has real filesystem access to its own
machine regardless.

### Decision provenance: `snapshot_ref`

Every ledger entry -- fresh judgment or [ledger cache
hit](#the-cache-hit-exception) -- now carries `snapshot_ref`: the commit
SHA of `PROJECT_ROOT` at the moment that line was written, or `null` when
`PROJECT_ROOT` isn't a git repo (or `git` isn't installed, or the call
times out) -- this never fails the underlying `recommend_component` call.
Entries written before this shipped read back with `snapshot_ref: null`.

A cache-hit entry's `snapshot_ref` reflects the codebase state *when that
cache-hit line was written*, not the original judgment's -- to see the
original judgment's snapshot, look up the entry named in its
`ledger_entry_id`/`original_verdict_timestamp` fields instead.

[`export_ledger_provenance`](#tool-export_ledger_provenance) packages one
entry's full record -- checklist, candidates, verdict, `snapshot_ref` --
into a markdown block you can paste into a PR or issue by hand.
[`post_ledger_provenance_to_github`](#tool-post_ledger_provenance_to_github)
posts that same artifact automatically, idempotently, using a personal
`GITHUB_TOKEN` rather than a GitHub App (see that tool's docs for why).
[`backfill_ledger_snapshot_ref`](#tool-backfill_ledger_snapshot_ref)
reconstructs a best-effort `snapshot_ref` for entries that predate the
field, always clearly labeled as reconstructed wherever it's rendered.

### Referential integrity: `file_path` / `live_status`

`recommend_component` optionally accepts `file_path` (see [Tool:
`recommend_component`](#tool-recommend_component)) -- usually not known at
call time, since the decision typically precedes the file existing. When
set, [`check_ledger_liveness`](#tool-check_ledger_liveness) can later
check whether that file still exists and still mentions
`chosen_candidate`:

- **`live`** -- the file exists and mentions `chosen_candidate`.
- **`orphaned`** -- `file_path` is set but the file no longer exists.
- **`unknown`** -- no `file_path` was ever recorded, the path escapes
  `PROJECT_ROOT`, or the file exists but `chosen_candidate` can't be
  confirmed in it. Deliberately the default outcome for anything
  ambiguous: a false `"orphaned"` is worse than a lingering `"unknown"`.
- **`dangling`** -- only ever produced by
  [`sweep_ledger_liveness`](#tool-sweep_ledger_liveness), never by
  `check_ledger_liveness` on its own: a cluster of 2+ entries sharing a
  `feature_id` where none of them resolved to `"live"`. Graph-level
  analysis across a project's whole entry set, not a single-entry check
  -- see that tool's docs for why `feature_id` is the grouping used.

`check_ledger_liveness` remains on-demand and single-project;
[`sweep_ledger_liveness`](#tool-sweep_ledger_liveness) is the
scheduled/batch counterpart -- meant to be invoked by your own cron/CI,
since Pattern has no scheduler of its own. `live_status`/`last_verified_live`
start `"unknown"`/`null` on every entry at write time and only ever
advance via a `check_ledger_liveness`/`sweep_ledger_liveness` call;
results are stored append-only in `~/.pattern/ledger_liveness.jsonl`
(override with `PATTERN_LEDGER_LIVENESS_PATH`, same "append, never mutate
the source line, most recent record wins at read time" convention as
`outcome_proxies.jsonl`, see [Outcome proxies](#outcome-proxies)) and
layered onto `ledger.jsonl`'s own entries at read time -- the ledger line
itself is never rewritten.

## Per-project decision memory

Pattern stores confirmed decisions locally in:

```
~/.pattern/memory.json
```

You can change the location with:

```
PATTERN_MEMORY_PATH
```

The file is organized by project:

```json
{
  "my-booking-app": [
    {
      "component_need": "price breakdown with fees and taxes",
      "domain": "Airbnb-style rental marketplace",
      "action": "custom_built",
      "source": "custom",
      "timestamp": "2026-08-25T14:32:00.000Z",
      "time_saved_minutes": 25
    }
  ]
}
```

`time_saved_minutes` is omitted from an entry entirely when the calling
agent didn't provide one -- it's never backfilled or estimated by Pattern.

Each project keeps its 50 most recent decisions. Older entries are
removed as new ones are added.

Only decisions explicitly recorded through `record_component_decision`
are saved. Pattern does not automatically save recommendations.

If an agent ignores or changes a recommendation, nothing is recorded
unless the agent explicitly calls `record_component_decision` with what
it actually did.

The memory file is local plaintext. Pattern does not send it anywhere.

`component_need` and `domain` are stored in this file, so avoid putting
sensitive information in them. See [SECURITY.md](./SECURITY.md).

A failure to write the decision file is returned as an error from
`record_component_decision`.

**No caching, by design.** Project memory (this file, `memory.json`) does
not cache recommendations. A previous decision is only additional context
for a new judgment. This is unrelated to the
[judgment ledger](#per-project-judgment-ledger)'s bounded cache-hit
exception, which lives in a separate file (`ledger.jsonl`) and is always
flagged (`served_from_ledger: true`) when it happens — see
[Known limitations](#known-limitations) for more.

## Security and privacy

Pattern uses the Anthropic API to make its recommendations, sending your
component need and your registered design system's candidate data (names,
props, descriptions, Figma data). It does not use web search.

Local project memory and the local call log are stored on the machine
running Pattern. They are not sent anywhere by Pattern itself.

The one exception is telemetry, on by default -- see
[Telemetry](#telemetry) below for exactly what it sends and how to turn
it off.

Review [SECURITY.md](./SECURITY.md) before putting sensitive information
into fields such as `component_need`, `domain`, or project IDs.

## Telemetry

On by default, as of v0.9.0. To turn it off:

```
PATTERN_TELEMETRY=0
```

(`false` and `no` also work, case-insensitively. Before v0.9.0 this was
opt-in -- `PATTERN_TELEMETRY=1` to turn on. In practice that meant almost
no signal: installs happened, but almost nobody set the env var. This
release flips the default and adds a second, standard event stream, both
described below, while keeping the same "never sent" guarantees.)

**The one-time notice.** The first time you run this version of Pattern
-- whether it's a brand-new install or an upgrade from an older version
-- it prints a short notice to stderr explaining all of this and how to
opt out. It prints exactly once, ever (tracked by a marker file at
`~/.pattern/telemetry_notice_shown`), then never again, regardless of
whether you act on it. There's no interactive y/n prompt: Pattern's stdin
is the MCP JSON-RPC channel the client uses to talk to it, so blocking on
stdin for a keypress would fight the protocol handshake instead of
showing a dialog -- a stderr notice is the safe equivalent for a stdio
MCP server. The same first-run moment also surfaces the enforcement
boundary, with the same constraint handled the same way -- see
[Enforcement boundary: hook + CI gate](#enforcement-boundary-hook--ci-gate).

**Why it exists.** Three things about real usage can't be answered from
this repo alone: whether people actually come back and use Pattern on a
second or third project on their own; how often a BYO Anthropic key
actually hits a rate limit or runs out of credit in real sessions, not
just the one time that happened during manual testing (see
[Known limitations](#known-limitations)); and, previously, whether
`recommend_component` gets called at all after install -- opt-in
telemetry couldn't answer that last one because the people who'd opt in
are a biased, tiny sample of everyone who installs.

**What gets sent, when on -- two event streams, same distinct ID:**

1. Pattern's own events, sent directly to PostHog:
   - An anonymous, randomly generated install ID -- a UUID created once
     and stored at `~/.pattern/install_id` (overridable via
     `PATTERN_INSTALL_ID_PATH`), never derived from your machine,
     username, or any other identifying information. This is the distinct
     ID both event streams use, so an install counts as the same "user"
     in either.
   - A one-way SHA-256 hash of `project_id`, truncated to 16 hex characters
     -- never the raw `project_id` string. The hash lets Pattern count how
     many *distinct* projects one install has used, without ever seeing what
     those projects are named.
   - On every `recommend_component` call that reaches the API or the ledger
     cache-hit shortcut: `verdict`, `confidence`, `reason`,
     `ensemble_triggered`, `estimated_cost_usd`, and `served_from_ledger` --
     the same distilled shape already written to the
     [local call log](#local-call-log), not new information.
   - On a failed Anthropic API call specifically: the HTTP status code and a
     coarse classification (`rate_limit`, `insufficient_credit`, or `other`)
     -- never the request or response body.
   - On every invocation of the `pattern-mcp` binary, immediately at
     startup: a single `pattern_cli_started` event carrying only which
     mode it ran in (`server` -- the normal MCP-server start, or `init` --
     the [connect wizard](#connect-pattern-to-your-mcp-client)), whether
     stdin was a terminal, and a coarse `invocation` bucket (`none`, `init`,
     or `other` -- never the raw arguments), and whether the process was
     started by the `init` wizard's own health check. This
     exists to separate real executions from npm registry traffic that
     never runs the code at all (security scanners, mirrors) -- something
     neither `recommend_component` counts nor `@posthog/mcp`'s handshake
     event below can answer, since both require getting further than a
     bare `npx pattern-mcp` run.
   - When `init` finishes: a single `pattern_cli_init_completed` event with
     only coarse outcomes -- whether a client was connected, whether the
     wizard's startup self-test passed, whether an Anthropic key was set
     (`valid`/`unverified`/`invalid`/`missing`, never the key), whether a
     design system was registered, whether enforcement was set up, and how
     many of the five steps finished. No paths, names or keys.
   - When `recommend_component`, `extract_requirements` or
     `register_design_system` fails: a single `pattern_cli_tool_error` event
     with the tool name and a coarse `error_class` (`missing_api_key`,
     `api_rate_limit`, `api_insufficient_credit`, `api_error`,
     `truncated_output`, `session_cap`, `bad_arguments`, `file_not_found`,
     `figma`, or `other`) -- never the message, paths or arguments.
   - On process exit, as of v0.14.0: a single `pattern_cli_exited` event
     carrying only a coarse reason (`sigint`, `sigterm`,
     `uncaught_exception`, `unhandled_rejection`,
     `fatal_startup_error`, or `stdin_closed` -- the client disconnected) and, for the two exception cases, the thrown
     value's constructor name (e.g. `TypeError`) -- never the error
     message or stack trace. Paired with `pattern_cli_started` so a start
     with no matching MCP handshake is diagnosable as a crash instead of
     silent.
2. Standard MCP tool-call analytics, via
   [`@posthog/mcp`](https://posthog.com/docs/mcp-analytics): which tool
   was called, call duration, and success/failure, so unique installs and
   call counts per tool (including `recommend_component`) are visible the
   same way any other MCP server's usage would be. This SDK's defaults
   would otherwise also capture full tool call arguments, full response
   text, and the raw thrown error message on a failure -- Pattern
   explicitly strips all of that (`$mcp_parameters`, `$mcp_response`,
   `$mcp_intent`, `$mcp_error_message`) before anything leaves the process,
   via its `beforeSend` hook. What's left is the same shape as stream 1
   above: counts, timing, and coarse success/failure, never content.

**What never gets sent, telemetry on or off:** `component_need`,
`domain`, `framework`, `existing_stack`, `requirements_checked` evidence,
the raw `project_id`, a failed API call's request or response body, or
your Anthropic API key. This is a claim about both event streams above,
not just the first one.

**Where it goes.** Both streams go to Pattern's PostHog project via its
public, write-only project key (safe to ship in source -- it can send
events, it cannot read data back). Set `PATTERN_POSTHOG_KEY` /
`PATTERN_POSTHOG_HOST` to point at a different project, e.g. for
self-hosting.

**Turning it off:** `PATTERN_TELEMETRY=0` disables both streams entirely
-- when it's off, `instrument()` (stream 2) is never even called, and
stream 1's `capture()` calls are no-ops.

## Cost

Pattern uses the Anthropic API, so `recommend_component` has a cost.

A typical single pass costs about $0.06–$0.10 with Sonnet 5 at current
pricing. Skip-listed primitives cost $0 because they're handled locally
and never reach the API. A [ledger cache hit](#the-cache-hit-exception)
also costs $0, for the same reason -- no API call happens.

### The `_meta` field

Every `recommend_component` and `extract_requirements` response includes
an internal `_meta` block reporting what that call actually spent. This
is not shown to the user automatically -- the calling agent has to
surface it, the same way it's separately instructed to show
`install_command` before running it (see
[above](#installation-commands-are-not-trusted)). Both tool descriptions
say so explicitly: surface `_meta.estimated_cost_usd` after the call,
since it's real spend against the user's own API key, not internal
bookkeeping.

```json
{
  "total_ms": 41516,
  "breakdown_ms": { "extract": 5006, "search": 3114, "score": 33396 },
  "tokens_used": { "input": 8400, "output": 620 },
  "estimated_cost_usd": 0.14,
  "scoring_fetch": { "attempted": true, "succeeded": true, "url": "https://ui.shadcn.com/docs/components/..." }
}
```

- `total_ms` -- wall-clock time for the call.
- `tokens_used` -- total input tokens (fresh + cache write + cache read,
  summed) and output tokens, read directly from the API response's own
  usage data.
- `estimated_cost_usd` -- computed from `tokens_used` at Pattern's
  configured model's current per-token rate (checked against Anthropic's
  pricing, not assumed). This is an estimate: it doesn't account for
  pricing changes Pattern hasn't been updated for, or any account-specific
  discounts.
- `breakdown_ms` -- how `total_ms` splits across `recommend_component`'s
  three internal phases.
- `scoring_fetch` -- whether step 4's single candidate-verification fetch
  (see [Fetch-grounded scoring](#search-limits-removed-in-018)
  below) actually happened for this response. `url` is `null` when
  `attempted` is `false` (no real candidate to verify, e.g. `reason:
  "no_candidates_found"` or `"skip_list"`). This is a diagnostic only --
  Pattern never uses it to auto-correct `requirements_checked` after the
  fact, since there's no safe fallback value for an unverified met/not-met
  call the way there is for a reference URL.

**How `breakdown_ms` is measured, and its one real caveat.** The bundled
call runs extraction and scoring inside a single model turn, with no
tools, so there is no natural place for separate stopwatches. `_meta.breakdown_ms`
reports `extract` (time until scoring starts), `search` (always `0` now that
search is gone), and `score`. Before 0.18, `extract` ended when the first
search call started; without search, the split is approximate.

**When the ensemble triggers** (see below), `_meta` reports the sum
across all reruns that actually happened -- total tokens and cost spent,
not the wall-clock time you waited. The three ensemble passes run with the
2nd and 3rd concurrent, so perceived latency is closer to ~2x one pass,
not the ~3x `total_ms` will show. Cost and token spend are genuinely
additive across reruns, which is what `_meta` is reporting there.
`scoring_fetch` is the one exception -- it isn't summed (a fetch either
happened for the specific pass whose evidence became the returned
`requirements_checked`, or it didn't), so it reports that winning pass's
own value, not an aggregate across all three.

Three things help keep the cost down without changing the decision process.

### Prompt caching

Pattern caches its system instructions using `cache_control: ephemeral`.

The instructions are the same across calls, so repeated requests don't
pay the full input cost for that block.

### Measured cache and fetch behavior (historical, pre-0.18)

`_meta.tokens_used.input_breakdown` splits input tokens into `fresh`,
`cache_write`, and `cache_read` (see [The `_meta`
field](#the-_meta-field)) -- added specifically to check assumptions
about caching against real numbers rather than guessing. Two real
findings so far:

- **A single, non-repeat call is not "all fresh."** The working
  assumption had been that only exact-repeat calls (the [ledger cache
  hit](#the-cache-hit-exception)) benefit from caching at all. A live
  test disproved that: a fresh, non-repeat toast-component call came back
  with roughly half its input tokens served from `cache_read`. A
  follow-up 4-case sample (2026-09-02, spanning a clean `use_existing`
  call, a `custom_build` call, and two historically boundary/inconsistent
  cases) confirmed this wasn't a fluke -- `cache_read` share stayed in a
  46-63% band across all four, regardless of call shape.
- **`fresh` (fully-priced, never-cached) tokens are driven by whether the
  call reaches step 6's Mobbin/Figma reference search, not by general
  complexity or the boundary-risk ensemble firing.** In that same
  4-case sample, the two `use_existing` calls had negligible `fresh`
  tokens (0.2%); both `custom_build` calls (which searched Mobbin/Figma)
  had 24-27.5% `fresh` -- even though, in both of those cases, the actual
  Mobbin *fetch* failed (`url_not_accessible`, 0 bytes returned). That
  rules out fetched-page content size as the driver for this cost --
  it's the extra Mobbin/Figma-restricted *search* calls themselves. This
  is why [`PATTERN_FETCH_MAX_CONTENT_TOKENS`](#search-limits-removed-in-018)
  was trimmed (a fetch-content cap can't fix a search-call cost) rather
  than split per-step as originally considered, and why reducing
  Mobbin/Figma search overhead is tracked as its own, differently-scoped
  future item rather than folded into that change.

### Search limits (removed in 0.18)

`PATTERN_SEARCH_BUDGET` and `PATTERN_FETCH_MAX_CONTENT_TOKENS` no longer
exist. Pattern makes no web search or fetch calls, so there is nothing to
cap. Older results in this README that mention search budgets, reference
verification, or `_meta.scoring_fetch` describe the pre-0.18 behavior.

### Choosing a cheaper model

You can change the model with:

```
PATTERN_MODEL
```

It defaults to:

```
claude-sonnet-5
```

You could use a cheaper model such as Haiku 4.5 without changing the code.

Before using a cheaper model in production, run the five validation cases
and compare its results with Sonnet's:

- Price breakdown
- Cancellation policy
- Earnings dashboard
- Image gallery
- Messaging inbox

The cheaper model hasn't been validated yet, so these results should be
treated as an open question rather than an established performance claim.

### Ensemble cost (boundary-risk cases only)

Pattern uses extra model calls only when a result is close enough to a
decision threshold that a small change in judgment could change the
verdict.

The default checklist has eight items, so coverage can only land on
these values:

```
0%
12.5%
25%
37.5%
50%
62.5%
75%
87.5%
100%
```

The decision thresholds are 40% and 80%.

That means results at 37.5%, 50%, 75%, and 87.5% are the cases where
changing the judgment on one requirement can flip the verdict.

For those cases, Pattern runs the full judgment three times and takes
the majority result.

For example:

```json
{
  "ensemble": {
    "triggered": true,
    "runs": ["use_existing", "custom_build", "use_existing"],
    "agreement": "2/3"
  }
}
```

If all three runs agree, the majority verdict is returned normally.

If they split 2/3, Pattern sets confidence to `"low"`. The disagreement
is surfaced rather than hidden.

Results at 0, 12.5, 25, 62.5, and 100% stay single-pass because one
changed requirement can't move them across either threshold.

**Other checklist sizes.** The same rule applies to any size, e.g. a
checklist you pass in or hand-edit: a result is a boundary case only when one
requirement flipping (met ± 1) would move coverage across 40% or 80%. For 8
items that is exactly the set above; for 9 or 10 items it is 3, 4, 7 or 8 met.
(Before, any checklist that was not exactly eight items always paid for the
extra passes.)

### Variance report (`stability`)

Whenever the ensemble runs two or more scoring passes, the result also carries
a `stability` block built from passes that were already paid for (no extra
call):

```json
"stability": {
  "passes": 2,
  "met_per_pass": [7, 7],
  "coverage_spread": "7 of 9 in every pass",
  "items_comparable": true,
  "split_items": [ { "requirement": "...", "met_votes": "1/2" } ],
  "unanimous_items": 8
}
```

`split_items` are the checklist items the passes disagreed on -- the ones to
look at by hand. With a checklist you pass in, items are matched by position
(the model sometimes shortens the wording); with an extracted checklist each
pass writes its own items, so `items_comparable` is `false` and only the
verdict agreement is reported. It is reporting only: it never changes the
verdict, coverage or confidence, and a single-pass result has no `stability`
block, because one pass carries no variance information.

### Scoring effort

The design-system scoring pass runs with `effort: medium` by default. On two
measured needs that was about 25-35% cheaper and about 2x faster than the
model's default effort (`high`) with the same verdict; item-level calls still
vary a little between runs at every setting, and there are no ground-truth
labels for that, so accuracy was not separately measured. Override with
`PATTERN_SCORE_EFFORT=low|medium|high|xhigh|max`. `PATTERN_SCORE_THINKING=disabled`
turns thinking off entirely (accepted on `claude-sonnet-5` only; it was the
cheapest and most stable setting but also the most lenient scorer). Thinking
tokens count against `max_tokens`, so the ceiling is 16384 on this path.

### Measured ensemble cost

The ensemble doesn't mean every call costs 3x.

In the latest five-case validation, Pattern made 15 outer calls:

- 8 stayed single-pass
- 7 triggered the ensemble
- 21 model passes were used for those 7 ensemble calls
- 29 total model calls across the test

That works out to about a 1.9x average multiplier across that test set.

The worst case is still 3x for an individual call when the ensemble is
triggered.

### What the ensemble can and cannot solve

The ensemble reduces the chance that one unlucky model judgment
determines the result. It doesn't eliminate uncertainty.

If the underlying evidence is genuinely ambiguous, three runs can still
disagree.

For example, the image-gallery validation case continued to flip between
outer runs. When that happened, the ensemble consistently reported a 2/3
split with `confidence: "low"`.

That's expected behavior: the tool is exposing uncertainty instead of
presenting an ambiguous result as certain.

### Session call cap

Pattern limits the number of API calls to 40 per server process by
default.

You can change this with:

```
PATTERN_SESSION_CAP
```

The cap protects against runaway agents, such as an agent stuck in a
retry loop or repeatedly asking for the same recommendation.

The 40-call default is based on the project's validation work. A
realistic project with roughly 25 components would use about 25 calls
for a full pass, leaving room for iteration.

Skip-listed primitives don't count because they never reach the API.

The counter lives in memory and resets when the server restarts.

If 40 calls is too low for your project, increase `PATTERN_SESSION_CAP`
rather than repeatedly restarting the server.

## Local call log

Every API call is recorded in a local log.

By default:

```
~/.pattern/calls.log
```

You can change the location with:

```
PATTERN_LOG_PATH
```

The log is local. Pattern does not send it anywhere.

Each API call adds one JSON line, for example:

```json
{
  "timestamp": "2026-08-24T21:12:43.882Z",
  "component_need": "cancellation policy display",
  "domain": "Airbnb-style rental marketplace",
  "framework": "React + Tailwind",
  "verdict": "custom_build",
  "confidence": "high",
  "reason": "scored",
  "coverage": "2/8 (25%)",
  "ensemble_triggered": false,
  "checklist_source": "extracted",
  "total_ms": 44834,
  "estimated_cost_usd": 0.15
}
```

Additional fields appear when relevant:

- `ensemble_agreement` appears when the ensemble runs.

`checklist_source`, `total_ms`, and `estimated_cost_usd` mirror the
call's `_meta` block (see [Cost](#cost)) -- `total_ms` and
`estimated_cost_usd` are the same aggregated-across-reruns numbers when
the ensemble triggers, not per-pass figures.

The log deliberately does not contain:

- The full `requirements_checked` evidence
- Your Anthropic API key

It does contain `component_need` and `domain`, so avoid putting sensitive
information in those fields. See [SECURITY.md](./SECURITY.md).

The log directory is created automatically.

If Pattern cannot write to the log because of permissions, a read-only
filesystem, or a full disk, it reports the problem to stderr but does not
fail the tool call.

### Review a log

You can summarize a log with:

```
node summarize-log.js [path]
```

If no path is provided, it uses the same default location as the server.

The summary includes:

- Verdict and confidence breakdown
- Reason breakdown
- Ensemble trigger and agreement rates
- Reference-source grounding rates for custom builds
- Component needs that were requested more than once

Repeated component needs can be useful to investigate alongside the
[session call cap](#session-call-cap).

## Known limitations

### Model judgment can vary

Pattern's search results can stay the same while the model's
interpretation of those results changes between runs.

Validation found cases where two runs found the same named components
using the same search queries but judged the same evidence differently.

For example, the model interpreted an Export action as present in one
run and absent in another.

This is a limitation of model-based evidence judgment, not necessarily a
search or code problem.

The boundary-risk ensemble exists to detect and surface this uncertainty.

### A missed match in your own design system isn't automatically re-checked

In [design-system mode](#tool-register_design_system), the model can say
`custom_build`/`no_candidates_found` even when a real, relevant candidate
is sitting right in its own prompt -- a reading-comprehension miss over a
fully-known candidate list, not a live-search gap. The boundary-risk
ensemble above doesn't catch this: it only re-checks a `"scored"` result
near the 40%/80% threshold, never a `"no_candidates_found"` verdict.

The [keyword-overlap safety net](#a-safety-net-for-a-missed-match) flags
this risk (`design_system_recall_check` on the response) but does not fix
it -- it's a detection layer, not a re-check. The actual fix (re-running
the model on a suspicious `no_candidates_found` verdict, the same way a
close `"scored"` call already gets re-checked) is scoped but not built.
`custom_build` ledger entries in this mode also record zero candidates
(`candidates_evaluated: []`), same as the external-library path, so a
genuine miss and a correct "nothing here" still look identical in
`read_ledger` afterward unless the recall check happened to catch it.

### A staged pipeline was evaluated and not adopted

To address the variance above, an alternative architecture was built and
tested: splitting the single bundled judgment call into separate stages
(extract requirements, search evidence, score coverage), on the theory
that isolating each step would make results more consistent and easier
to diagnose.

A pilot comparison (5 cases, 3 repeated runs per case, per
architecture) found no consistent benefit. The staged pipeline improved
consistency on one boundary-risk case but was less consistent than the
bundled pipeline on another, including one run that failed outright.
Net accuracy against hand-graded gold answers was statistically
indistinguishable between the two architectures, and the staged
pipeline cost roughly **2x** the bundled pipeline's call volume across
the board, not only on the boundary-risk cases it was expected to help
most.

Pattern ships the bundled pipeline. The staged implementation remains
in the repo (`src/staged/`) as an evaluated, unshipped experiment, not
a supported alternative.

**`extract_requirements` is not a revival of this.** It's a standalone
tool for inspecting the extraction step's output before an agent commits
to `recommend_component`'s search+score budget -- an opt-in visibility
tool, not an internal re-architecture. `recommend_component`'s own
pipeline is still fully bundled; nothing about this evaluation changed.

### No caching, by design

Every recommendation searches and scores again -- with one bounded
exception (see below).

This means a recommendation can change as component libraries change.
For example, a later shadcn/ui release can introduce a component that
changes a previous `custom_build` result.

Do not build a second, unbounded cache of recommendations at the
calling-agent layer on top of Pattern's own. If you add caching there,
keep it session-scoped.

[Project decision memory](#per-project-decision-memory) does not change
this. It provides context from previous decisions, but every
`recommend_component` call still performs a fresh search and scoring
pass.

The one deliberate exception is the
[judgment ledger's cache-hit path](#the-cache-hit-exception): a later
call matching an exact, recent, high-confidence prior judgment can be
served without a fresh search+score. It's bounded (exact
component_need/domain/framework/conventions match, a staleness TTL) and
always self-identifies via `served_from_ledger: true` and
`reason: "ledger_cache_hit"` -- so a calling agent that wants a guaranteed
fresh check on every call should look for that flag and treat it the same
as any other verdict it wants to double-check.

### The skip-list is still evolving

The primitive skip-list is a starting point and has not yet been
validated against broad real-world usage.

Watch for two failure modes:

- Agents calling Pattern for things that should have been skipped.
- Agents building generic UI for something that should have been on the
  skip-list.

The local call log can help identify both patterns.

### Pattern needs internet access

Pattern requires outbound access to:

```
api.anthropic.com
```

If `TYPESAFE_API_KEY` is set (or `PATTERN_SCORER=jev`), candidate evidence
(component names, doc comments and summaries, which may reflect real product
or UI text) and the component need are also sent to `api.typesafe.ai`. With
no TypeSafe key and `PATTERN_SCORER` unset, nothing goes to TypeSafe; set
`PATTERN_SCORER=anthropic` to keep it off even when the key is present.

It will not work in an environment that blocks general outbound internet
access.

### Requirements and coverage are judgment calls

Requirement extraction and evidence scoring are performed by the model.

Pattern adds safeguards such as:

- Structured requirements
- Server-side coverage recalculation
- Decision thresholds
- Boundary-risk ensembling
- Grounding checks for reference URLs

But the underlying interpretation of whether evidence satisfies a
requirement is still model judgment.

When introducing Pattern into a new workflow, spot-check early results
against the actual components before relying on it unattended.

## Setup preflight

`npx -p pattern-mcp pattern doctor` checks the things that usually go wrong before the first run: which MCP config holds the server, which keys are in that server's own `env` block (a shell export or project `.env` is never read), whether the Figma token is well-formed (`--online` also tests it), the resolved project root and id, and whether a design system is registered. It never prints a secret.

`npx -p pattern-mcp pattern fetch-figma <file_key>` downloads a Figma file with your token and saves it to `.pattern/figma/<key>.json` only if it is a real file (never an error body). Register it with `figma_json_path`. For big files, scope with `figma_pages`.

TDQS

A4.5/5.0

Scored across 1 tool

Disambiguation5/5

With only a single tool, there is no possibility of confusion between tools. The tool's purpose is clearly defined and unambiguous.

Naming Consistency5/5

The tool name 'recommend_component' follows a clear verb_noun pattern, which is consistent with common naming conventions. No other tools exist to introduce inconsistency.

Tool Count3/5

A single tool is on the thin side for most servers, but it serves a specific, well-defined purpose. It is not trivial and may be sufficient for the intended scope, though additional related tools could add value.

Completeness5/5

The tool fully covers its stated purpose: making a judgment between using existing components or building custom. It provides a structured verdict without missing essential aspects like referencing Mobbin or shadcn/ui. No obvious gaps within the defined domain.

Maintenance

ActivityActive
ResponsivenessNo issues