Skip to main content
Glama
J-X0
by J-X0
README.md
# westmere-recsys

Recommendation engine for Meridian Commercial (Project Westmere). It ranks
retail merchandising candidates with a Thompson-sampling bandit, and it does so
under a hard operating constraint: **spend is capped per tenant per calendar
month**.

The delivery is an MCP server: JSON-RPC 2.0 over newline-delimited stdio,
implemented on the standard library (no `mcp` SDK required), exposing three
tools: `recommend`, `record_feedback`, `budget_status`.

## How the spend cap shapes the design

Every contextual score comes from a `ScoringProvider` (an LLM call in
production). That call is the operating spend. The engine treats exploration as
something it has to pay for:

- Before scoring an arm, the engine tries to charge the provider's per-call
  cost against the tenant's monthly budget via an atomic check-and-charge
  ledger (`SpendLedger.try_charge`).
- If the charge succeeds, it pays for a contextual prior and blends it with the
  Thompson sample. This is exploration.
- If the charge is refused because the cap is reached, the engine serves that
  arm from the bandit's posterior mean at zero cost. This is exploit-only
  degradation.

Consequences, all covered by tests:

- Recorded spend can never cross the cap (`try_charge` is atomic; no partial
  charges).
- The service keeps returning recommendations after the budget is gone.
- Every decision emits an `AuditRecord` (spend before/after, provider calls,
  chosen arms, which were explored, and what constraints rejected which arms).

## Architecture

Requests flow through four layers, each with one job:

```
JSON-RPC stdio  ->  WestmereService   ->  RecommenderEngine  ->  ThompsonBandit
  (server.py)       validate/time/log     rank under budget       posterior draws
                                             |         |
                                       SpendLedger  ConstraintFilter
                                       (the cap)    (hard rejections)
                                             |
                                       ScoringProvider (stub | real)
```

- `server.py` is transport only: parse a line, dispatch, write a line. Its
  `dispatch` function is pure so it can be tested without any I/O.
- `WestmereService` is the boundary: nothing reaches the engine without being
  validated. It owns config wiring, timing, and structured logging.
- `RecommenderEngine` is the algorithm: constraint filter, then per-arm
  charge-and-score, then rank. It never talks to the transport or the config.
- `SpendLedger`, `ConstraintFilter`, `ThompsonBandit` and `ScoringProvider` are
  independent collaborators, each unit-tested in isolation.

Design decisions that had a real alternative are recorded in `docs/adr/`.

## Layout

```
westmererecsys/
  domain.py        Arm, Context, Recommendation, AuditRecord, RecommendationResult
  bandit.py        ThompsonBandit (Beta-Bernoulli posterior, seedable RNG)
  budget.py        SpendLedger, per-tenant/per-month cap, BudgetExceeded
  constraints.py   ConstraintFilter + builtin constraints
  engine.py        RecommenderEngine: ties the above together
  config.py        Config: load/validate from dict, JSON file, or env
  errors.py        WestmereError / ConfigError / ValidationError
  logging_setup.py JSON-per-line structured logging to stderr
  service.py       WestmereService: input validation, timing, JSON results
  server.py        MCP JSON-RPC stdio entry point (dispatch + serve loop)
  providers/
    base.py        ScoringProvider interface (defines cost_cents)
    stub.py        StubScoringProvider: deterministic, offline
    real.py        RealScoringProvider: OpenAI-compatible, credential-gated
tests/
```

## Setup and tests

```
make install     # create .venv and install the package with dev extras
make test        # run the suite (offline, no API key needed)
make lint        # byte-compile the package and tests
```

`PY` overrides the interpreter, e.g. `make test PY=python3`.

The entire suite runs on the deterministic stub provider with no network
access and no API key.

## Running the server

```
make run                                  # cap 1000c/tenant/month, stub provider
python3 -m westmererecsys.server --config config.example.json --catalogue catalogue.example.json
python3 -m westmererecsys.server --default-cap-cents 500 --log-level DEBUG
```

Flags: `--config` (JSON config file), `--catalogue` (JSON array of items used
when a `recommend` call omits `candidates`), `--default-cap-cents` (used when no
`--config` is given), `--log-level`. Environment overrides `WESTMERE_*` (e.g.
`WESTMERE_DEFAULT_CAP_CENTS`, `WESTMERE_PROVIDER`, `WESTMERE_SEED`) are applied
on top of the file. A missing or invalid config fails fast with exit code 2.

It speaks JSON-RPC 2.0, one message per line, on stdin/stdout. Example session:

```
{"jsonrpc":"2.0","id":1,"method":"initialize"}
{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"recommend","arguments":{"tenant_id":"store-42","candidates":[{"id":"a"},{"id":"b"}],"k":1}}}
```

### Config file keys

`default_cap_cents` (required), `per_tenant_caps`, `provider` (`stub`|`real`),
`provider_cost_cents`, `explore_weight`, `seed`, `prior_alpha`, `prior_beta`,
`require_in_stock`, `exclude_categories`, `max_price_cents`, `max_candidates`,
`max_k`, `log_level`.

## Failure handling and degradation

- Bad request input (missing `tenant_id`, `k` out of range, malformed
  candidate, too many candidates) raises `ValidationError`, surfaced to the
  client as an MCP tool error (`isError: true`) rather than crashing the loop.
- Missing/invalid config or catalogue files raise `ConfigError` at startup.
- `max_candidates` / `max_k` bound per-request work so one call cannot exhaust
  resources.
- If the provider raises mid-request, the engine refunds that arm's charge (the
  tenant is not billed for spend that produced nothing) and degrades the arm to
  posterior-mean scoring. The `provider_errors` count is in every audit record.
- Once the monthly cap is reached, the engine serves exploit-only at zero cost
  and logs a warning; the service keeps responding.

## Using the engine

```python
from westmererecsys import (
    Arm, Context, RecommenderEngine, SpendLedger, ThompsonBandit,
    ConstraintFilter, require_in_stock,
)
from westmererecsys.providers.stub import StubScoringProvider

engine = RecommenderEngine(
    provider=StubScoringProvider(cost_cents=2),
    ledger=SpendLedger(default_cap_cents=500, per_tenant_caps={"vip": 5000}),
    bandit=ThompsonBandit(seed=7),
    constraints=ConstraintFilter([require_in_stock()]),
)

ctx = Context(tenant_id="store-42", features={"segment": "loyal"})
catalogue = [Arm(f"sku-{i}", f"Item {i}", "apparel") for i in range(20)]

result = engine.recommend(ctx, catalogue, k=5)
for rec in result.recommendations:
    print(rec.arm_id, round(rec.score, 3), "explored" if rec.explored else "exploit")

# Fold observed outcomes back in (reward in [0, 1]):
engine.record_feedback("sku-3", reward=1.0)
```

## Production provider

`RealScoringProvider` calls an OpenAI-compatible chat endpoint. It reads
`WESTMERE_LLM_API_KEY` and optional `WESTMERE_LLM_BASE_URL`, and it raises
`ProviderUnavailable` rather than silently degrading when credentials are
missing, `httpx` is not installed, or the response cannot be parsed. Install
its dependency with `pip install .[real]`.

## Known limitations

- **State is in-memory.** Bandit posteriors and the spend ledger live in the
  server process. Restarting resets learning and, more importantly, resets
  spend to zero for the month. `ThompsonBandit.export_state`/`load_state` exist
  as the persistence seam, but the ledger has no durable store yet. A restart
  loop could let a tenant exceed the intended monthly cap. See
  `docs/adr/0005-in-memory-state.md`.
- **Single process, no concurrency control.** The stdio loop handles one
  request at a time. `SpendLedger.try_charge` is atomic within a process but
  there is no cross-process locking, so running multiple servers against one
  tenant would not share a budget.
- **Month boundaries use the server clock in UTC.** A tenant billed in another
  timezone may see the cap reset a few hours early or late.
- **The blend weight is static.** `explore_weight` is a fixed constant, not
  tuned per tenant or decayed over time.
- **`RealScoringProvider` is not covered by the test suite** (it needs live
  credentials). Its parsing and credential-guard logic are tested; the network
  call itself is not.

---

*Meridian Commercial is an illustrative client; this repository is a self-directed reference implementation built to work end to end.*

TDQS

B3.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: recommend generates candidate rankings under budget, record_feedback logs observed rewards, and budget_status reports spend/remaining cap. There is no meaningful overlap between them.

Naming Consistency3/5

The names are readable but not fully consistent: 'record_feedback' follows a verb_noun pattern, 'recommend' is a bare verb, and 'budget_status' is noun_noun without an action verb. A consistent set like 'recommend_items', 'record_feedback', and 'get_budget_status' would improve predictability.

Tool Count5/5

Three tools is a well-scoped size for a focused recommendation/bandit service. Each tool covers a necessary part of the core workflow: recommending, recording feedback, and checking budget.

Completeness4/5

The core loop of recommend -> record_feedback -> check budget is covered, and there are no dead ends in that workflow. However, there is no tool for managing tenants, candidate items, or budget configuration, which are minor gaps if the server is expected to handle those resources.

Maintenance

ActivityInactive
ResponsivenessNo issues