Skip to main content
Glama
README.md
# woollama

**Web Over Ollama (and Llamas).** An MCP + OpenAI router for AI desktops.

πŸ“– **Documentation: [woollama.readthedocs.io](https://woollama.readthedocs.io/)**

woollama sits between AI clients (Cursor, the OpenAI SDK, Claude Desktop,
cosmic-fabric, anything that speaks OpenAI or MCP) and AI backends (Ollama,
Anthropic, fabric, lackpy, filesystem MCPs, anything that speaks OpenAI or
MCP). It composes them into orchestrated calls without inventing a new
protocol.

```
                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                          β”‚   AI clients        β”‚
                          β”‚   (any OpenAI or    β”‚
                          β”‚    MCP client)      β”‚
                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚            woollama                  β”‚
                  β”‚  OpenAI server  +  MCP server        β”‚
                  β”‚  ───────────────────────────────     β”‚
                  β”‚  routes models, tools, executors     β”‚
                  β”‚  composes patterns + tools + models  β”‚
                  β”‚  into named recipes                  β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚                                      β”‚
              β”Œβ”€β”€β”€β”΄β”€β”€β”€β”€β”                            β”Œβ”€β”€β”€β”€β”΄β”€β”€β”€β”€β”
              β”‚ MCP    β”‚  tools, prompts, resources β”‚ OpenAI  β”‚  inference
              β”‚ tool   β”‚                            β”‚ compat  β”‚
              β”‚ serversβ”‚                            β”‚ backendsβ”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜                            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              fabric-mcp, lackpy,                   Ollama, Anthropic,
              filesystem, git, …                    vLLM, llama.cpp, …
```

## Status

**The Rust daemon `woollamad` β€” a multi-backend router, both surfaces live,
published to crates.io + PyPI.** woollama works end-to-end as:

- an **OpenAI-compatible server**: `/v1/chat/completions` (pass-through *and*
  hidden chat-loop orchestration of recipes, both with `stream:true` β†’ OpenAI
  SSE), `/v1/models`, `/v1/tools`, and a **stateful surface** β€”
  `/v1/responses` + `/v1/conversations` (OpenAI Responses/Conversations shape;
  see below);
- an **MCP server** to its own clients β€” over **stdio** (`woollamad mcp`) and
  over **Streamable HTTP** at `/mcp`, mounted on the *same port* as `/v1/*`. It
  re-exports every discovered downstream tool (namespaced, with `output_schema`)
  plus a `chat` verb that emits live tool-progress notifications β€” i.e. it's an
  MCP aggregator.

It routes inference across **multiple backends** by `<provider>/<model>` β€”
`ollama` (local), `anthropic`, `openai`, `groq`, `together`, `openrouter`, and
**any OpenAI-compatible endpoint** you add in `inferencers.toml` (e.g.
self-hosted vLLM) β€” plus `claude-code/<model>`, a keyless path to Claude via the
local CLI (tool-less, or as an **executor** that runs a recipe's allow-listed
MCP tools itself β€” tool delegation). Config is file-driven (`mcp.json`,
`recipes.toml`, `inferencers.toml`).

**Stateful conversations** route *handles*; backends own the *state* β€” woollama
never stores transcripts in its own system. Two state-owning backends:
`claude-resume` (`claude --resume`, for `claude-code` models; keyless, the Claude
session owns the bytes) and `managed-agents` (Anthropic's Managed Agents, for
`claude-agent` models; `ANTHROPIC_API_KEY`, Anthropic hosts the session β€” and
exposes the transcript, so `/v1/conversations/{id}/items` works). Models with no
state-owning backend (ollama/cloud/recipe) are stateless β€” the caller owns
history (`store:false`). Long-lived MCP
connections. Served on **both a Unix socket** (`$XDG_RUNTIME_DIR/woollama.sock`,
mode 0600 β€” the default for local MCP clients) and an ephemeral loopback TCP
port; never `0.0.0.0` without explicit opt-in.

**Current status and what's next live in
[`docs/roadmap.md`](docs/roadmap.md).**

> **The Rust port is done (v0.5.x).** `woollamad` is the canonical router,
> published to crates.io (`cargo install woollama-server`) and PyPI
> (`pip install woollama`). The Python in `src/woollama/` is kept as the
> reference server and differential-test oracle β€” not deleted. See
> [`docs/rust-transition.md`](docs/rust-transition.md) for the (completed)
> transition criteria.

See `docs/architecture.md` for the full target design and
`docs/build-log.md` for the slice-by-slice history.

## Quick taste

The router is OpenAI-compatible, so any OpenAI client can drive it:

```python
import openai
c = openai.OpenAI(base_url="http://127.0.0.1:<port>/v1", api_key="x")

# Pass-through to Ollama
r = c.chat.completions.create(
    model="ollama/qwen3:14b-iq4xs",
    messages=[{"role": "user", "content": "Hi"}],
)

# Orchestrated: a recipe (system prompt + tools + model), transparent to the
# client. The chat-loop happens inside woollama; client sees only the final answer.
r = c.chat.completions.create(
    model="woollama/streamer",
    messages=[{"role": "user", "content": "Please count to 4."}],
)

# A device-aware inferencer (declares `management_url`): "default" resolves to
# whatever's currently loaded, and woollama loads/queues on demand β€” no
# not-loaded error, no wedge.
r = c.chat.completions.create(
    model="device/default",
    messages=[{"role": "user", "content": "Hi"}],
)

# Pass-through also covers images and embeddings:
img = c.images.generate(model="device/Z-Turbo", prompt="a small blue teapot")
vec = c.embeddings.create(model="device/Embedder", input="hello world")
```

woollama serves on **two transports at once**: a Unix socket at
`$XDG_RUNTIME_DIR/woollama.sock` (mode 0600 β€” the default for local MCP clients,
since a connectable socket can spend the router's API keys) and an ephemeral
loopback TCP port written to `$XDG_RUNTIME_DIR/woollama.addr` for clients to
discover. The `<port>` above is that ephemeral port. Same pattern as a local
`fabric --serve` instance.

## Install

The router is **`woollamad`** β€” a small Rust daemon. The Python implementation is
kept as a reference server and the differential-test oracle (see below), but
`woollamad` is the canonical router.

**From crates.io** (once published β€” `cargo install` ships only the binary, so
bring your own `mcp.json`):

```sh
cargo install woollama-server     # installs the `woollamad` binary
woollamad                         # starts the router; prints its address
```

**From this checkout** (works today; includes the bundled example MCP servers):

```sh
git clone https://github.com/teaguesterling/woollama
cd woollama
cargo build --release             # builds target/release/woollamad
./target/release/woollamad        # starts the router; prints its address
```

On startup `woollamad` prints its `OpenAI base_url` (e.g.
`http://127.0.0.1:<port>/v1`) β€” copy that into your OpenAI client. (It's also
written to `$XDG_RUNTIME_DIR/woollama.addr` for programmatic discovery, and it
serves the same surface over the `woollama.sock` unix socket.)

### Python 3.14: you will build from source, slowly

The `woollama` **Python** package depends on `woollama-core`, a compiled extension. We ship
wheels for **CPython 3.11–3.13 only** β€” pyo3 could not build against 3.14 until recently, and the
cp314 wheels are not published yet (#43).

Our metadata says `requires-python = ">=3.11"` with no upper bound, which is true β€” 3.14 works β€”
but on 3.14 there is no wheel, so the install falls back to **building the extension from
source**. That needs a Rust toolchain and takes minutes rather than seconds.

The trap is that `uv` picks an interpreter for you. `uv sync` in a project that merely allows
3.11+ will happily select 3.14 and then do a source build, or fail outright if no Rust toolchain
is present β€” and the error names `python-source` and maturin, neither of which mentions us. If
you want wheels today, pin the interpreter:

```sh
uv sync --python 3.12            # or any of 3.11-3.13
```

Any project depending on `woollama-core` **below 0.9.0** should pin 3.11–3.13 outright: those
versions cannot build on 3.14 at all, and every `woollama-core` sdist at or below 0.8.1 is
unbuildable on *any* Python (fixed in 0.8.2 β€” see the changelog for v0.14.4).

### The Python reference server

The original Python implementation still runs and is used as the live oracle that
keeps `woollamad` honest:

```sh
uv sync                           # creates .venv and installs deps
uv run woollama                   # the Python reference server
```

> **Prerequisite for the examples below:** they use `ollama/qwen3:14b-iq4xs`, so
> install [Ollama](https://ollama.ai), `ollama serve`, and
> `ollama pull qwen3:14b-iq4xs`. **No Ollama?** Use the keyless Claude path
> instead β€” `model="claude-code/haiku"` (needs the `claude` CLI logged in) β€” or
> any cloud model with its key set (see [Configuration](docs/configuration.md)).

### Tests & lint

```sh
# Rust (woollamad): the daemon's own suites
cargo test --tests --features test-fixtures
cargo build --release            # so the live oracle can spawn the binary

# Python: hermetic suite + lint
uv run --extra dev pytest        # hermetic suite (live tests are opt-in: -m integration)
uv run ruff check .              # lint β€” the CI gate

# The live differential oracle β€” same tests, against woollamad by default:
uv run --extra dev pytest -m integration            # targets target/release/woollamad
WOOLLAMA_TEST_CMD="python -m woollama" \
  uv run --extra dev pytest -m integration          # opt in to the Python reference
```

CI (`.github/workflows/ci.yml`) runs the Rust + Python gates on every push to `main` and PR.
For the same lint gate locally on commit, opt into the pre-commit hook:

```sh
uv tool install pre-commit && pre-commit install
```

Lint only β€” the project does not use `ruff format` (lines are hand-wrapped,
`E501` is ignored), so there is no formatter step in either gate.

## Design principles

1. **Two standards, neither extended.** MCP for tool/prompt/resource
   discovery and execution; OpenAI chat-completions for the inference
   primitive. woollama is a router between them.
2. **Local-only, ephemeral by default.** Random loopback port, persisted
   address file for discovery, never `0.0.0.0` without explicit opt-in β€” and
   the opt-in requires an auth token (`WOOLLAMA_TOKEN`; woollama refuses to
   start off-loopback without one). The router holds API keys and routes to
   local resources β€” it should not be LAN-reachable unauthenticated.
3. **The model namespace is the universal addressing scheme.** Raw inferencers
   (`<provider>/<model>`, e.g. `ollama/X`, `anthropic/X`, `claude-code/X`) and
   full recipes (`woollama/<recipe>`) are all addressable through OpenAI's
   standard `model` field. No new wire format.
4. **woollama owns routing, not inference or tools.** It uses other people's
   inference engines (Ollama, Anthropic, …) and other people's tool servers
   (any MCP server β€” filesystem, git, lackpy, …). It composes them.
5. **she talks to llamas.**

## What works today

- OpenAI surface: `/v1/models`, `/v1/chat/completions` (pass-through +
  recipe orchestration, both with `stream:true` β†’ OpenAI SSE), plus
  `/v1/images/generations` and `/v1/embeddings` pass-through (see below).
  (`/v1/tools` introspection is Python-reference-only today β€” see
  [#23](https://github.com/teaguesterling/woollama/issues/23))
- **Stateful surface**: `/v1/responses` (stateless subset, incl. `stream:true` β†’
  OpenAI Responses SSE, + stateful) and `/v1/conversations` (create/list/get/
  delete, plus `items` where the backend exposes its transcript). woollama routes
  conversation *handles*; backends own state (woollama never stores transcripts
  itself) β€” `claude-resume` for `claude-code` models, `managed-agents` (Anthropic
  Managed Agents) for `claude-agent` models, with an interactive
  `requires_action` pause/answer path; models with no state-owning backend are
  stateless (`store:false`)
- Multi-backend routing by `<provider>/<model>`: ollama (incl. `num_ctx` honored
  via ollama's native `/api/chat`), anthropic, openai, groq, together,
  openrouter, `claude-code`, + any OpenAI-compatible endpoint via
  `inferencers.toml`
- **Tool delegation**: a `claude-code` recipe with tools runs as an *executor* β€”
  Claude owns the agentic loop and calls the recipe's allow-listed MCP tools
  itself (per-recipe `--mcp-config` + `--allowedTools` containment)
- **Image + embedding pass-through**: `POST /v1/images/generations` and
  `POST /v1/embeddings` forward a `<provider>/<model>` request to that
  inferencer's own OpenAI-compatible endpoints, mirroring the chat pass-through
  (non-streaming; unknown namespace β†’ `400`)
- **Model pooling / device-aware inferencers**: an inferencer that declares a
  `management_url` gets on-demand model loading (no more bare not-loaded
  errors), stable **virtual model names** (`<provider>/default`, config
  `virtual` aliases), and per-model request queuing with `503` + `Retry-After`
  backpressure instead of a wedge β€” plus queue-aware LRU eviction to fit
  `pool_max` loaded models. Fully additive: an inferencer with no
  `management_url` is unaffected. Ships in **both** `woollama` (Python) and
  `woollamad` (Rust); pooling currently covers `/v1/chat/completions` only
  (`/v1/responses` isn't pooled yet)
- MCP server side: stdio (`woollamad mcp`) **and** Streamable HTTP at `/mcp` on
  the same port β€” recipes as **parameterized prompts** (their `{{var}}` tokens β†’
  arguments), a `chat` verb (with live tool-progress notifications), and every
  downstream tool re-exported with its `output_schema` (aggregator)
- **MCP client side, stdio *and* HTTP**: an `mcp.json` server is either a
  `command` (stdio subprocess) or a `url` (Streamable HTTP). Since woollama also
  *serves* `/mcp`, the `url` form is how **one woollamad consumes another** β€” an
  inference-holding instance reaching a tools-only instance's namespace, so the
  tools instance can hold zero provider API keys. Credentials are `${VAR}` in
  `headers`, validated fail-closed at load. `woollamad check-config` validates
  every config file and exits non-zero on any problem
- **Pattern templating** on woollama's own `/w1/` namespace (not OpenAI's `/v1/`):
  parameterized recipes/patterns with `{{var}}` substitution β€”
  `GET /w1/patterns` (discovery), `POST /w1/patterns/{name}/render` (assemble),
  `POST /w1/patterns/{name}/run` (render + infer, streaming). Patterns also come
  from a fabric-style directory scan (`[patterns]`) and a **fabric backend**:
  woollama can run/own `fabric --serve`, surface its library on `/w1/`, and
  transparently proxy fabric's API at `/fabric/*`. Pattern backends are pluggable
  (the `PatternBackend` trait β€” see `docs/extending.md`)
- File-driven config (`mcp.json`, `recipes.toml`, `inferencers.toml`), multi-
  MCP-server discovery + unified tool registry, long-lived MCP connections
- Recipe allow-list enforced as a security boundary (in-loop AND in delegation);
  served on a **Unix socket + loopback TCP**, address discovery file; CI
  (ruff + hermetic suite, 3.11/3.12)

## Not yet (next on the roadmap)

- The live, interactive Claude-in-tmux session backend (a separate Rust session
  driver) β€” gated on spikes that need a real terminal. (The interactive
  `requires_action` path itself already works via the managed-agents backend.)
- cosmic-fabric actually consuming the conversations surface (the last open
  integration milestone). The generic `store-backed` mechanism + two reference
  store providers (MCP + REST) already ship; what's pending is the cross-repo
  wiring. (Pattern templating + the fabric backend it needed have **shipped** β€”
  see `docs/patterns.md`.)
- lackpy re-pinning to the now-published `woollama-core` wheel.

Full scorecard, ordering, and pending verifications:
**[`docs/roadmap.md`](docs/roadmap.md)**.

## Origin

woollama is the production-grade rewrite of an architecture co-designed
in [cosmic-fabric](https://github.com/teaguesterling/cosmic-fabric), which
remains a frontend (and will use woollama as its router engine). The design
docs that brought woollama here:

- `docs/architecture.md` β€” the model/tool/executor router design
- `docs/naming.md` β€” how we landed on this name

## License

MIT β€” see `LICENSE`.