woollama
README.md
# woollama
**Web Over Ollama (and Llamas).** An MCP + OpenAI router for AI desktops.
π **Documentation: [woollama.readthedocs.io](https://woollama.readthedocs.io/)**
woollama sits between AI clients (Cursor, the OpenAI SDK, Claude Desktop,
cosmic-fabric, anything that speaks OpenAI or MCP) and AI backends (Ollama,
Anthropic, fabric, lackpy, filesystem MCPs, anything that speaks OpenAI or
MCP). It composes them into orchestrated calls without inventing a new
protocol.
```
βββββββββββββββββββββββ
β AI clients β
β (any OpenAI or β
β MCP client) β
ββββββββββββ¬βββββββββββ
β
ββββββββββββββββββββ΄ββββββββββββββββββββ
β woollama β
β OpenAI server + MCP server β
β βββββββββββββββββββββββββββββββ β
β routes models, tools, executors β
β composes patterns + tools + models β
β into named recipes β
ββββββββββββββββββββ¬ββββββββββββββββββββ
β
ββββββββββββββββββββ΄ββββββββββββββββββββ
β β
βββββ΄βββββ ββββββ΄βββββ
β MCP β tools, prompts, resources β OpenAI β inference
β tool β β compat β
β serversβ β backendsβ
ββββββββββ βββββββββββ
fabric-mcp, lackpy, Ollama, Anthropic,
filesystem, git, β¦ vLLM, llama.cpp, β¦
```
## Status
**The Rust daemon `woollamad` β a multi-backend router, both surfaces live,
published to crates.io + PyPI.** woollama works end-to-end as:
- an **OpenAI-compatible server**: `/v1/chat/completions` (pass-through *and*
hidden chat-loop orchestration of recipes, both with `stream:true` β OpenAI
SSE), `/v1/models`, `/v1/tools`, and a **stateful surface** β
`/v1/responses` + `/v1/conversations` (OpenAI Responses/Conversations shape;
see below);
- an **MCP server** to its own clients β over **stdio** (`woollamad mcp`) and
over **Streamable HTTP** at `/mcp`, mounted on the *same port* as `/v1/*`. It
re-exports every discovered downstream tool (namespaced, with `output_schema`)
plus a `chat` verb that emits live tool-progress notifications β i.e. it's an
MCP aggregator.
It routes inference across **multiple backends** by `<provider>/<model>` β
`ollama` (local), `anthropic`, `openai`, `groq`, `together`, `openrouter`, and
**any OpenAI-compatible endpoint** you add in `inferencers.toml` (e.g.
self-hosted vLLM) β plus `claude-code/<model>`, a keyless path to Claude via the
local CLI (tool-less, or as an **executor** that runs a recipe's allow-listed
MCP tools itself β tool delegation). Config is file-driven (`mcp.json`,
`recipes.toml`, `inferencers.toml`).
**Stateful conversations** route *handles*; backends own the *state* β woollama
never stores transcripts in its own system. Two state-owning backends:
`claude-resume` (`claude --resume`, for `claude-code` models; keyless, the Claude
session owns the bytes) and `managed-agents` (Anthropic's Managed Agents, for
`claude-agent` models; `ANTHROPIC_API_KEY`, Anthropic hosts the session β and
exposes the transcript, so `/v1/conversations/{id}/items` works). Models with no
state-owning backend (ollama/cloud/recipe) are stateless β the caller owns
history (`store:false`). Long-lived MCP
connections. Served on **both a Unix socket** (`$XDG_RUNTIME_DIR/woollama.sock`,
mode 0600 β the default for local MCP clients) and an ephemeral loopback TCP
port; never `0.0.0.0` without explicit opt-in.
**Current status and what's next live in
[`docs/roadmap.md`](docs/roadmap.md).**
> **The Rust port is done (v0.5.x).** `woollamad` is the canonical router,
> published to crates.io (`cargo install woollama-server`) and PyPI
> (`pip install woollama`). The Python in `src/woollama/` is kept as the
> reference server and differential-test oracle β not deleted. See
> [`docs/rust-transition.md`](docs/rust-transition.md) for the (completed)
> transition criteria.
See `docs/architecture.md` for the full target design and
`docs/build-log.md` for the slice-by-slice history.
## Quick taste
The router is OpenAI-compatible, so any OpenAI client can drive it:
```python
import openai
c = openai.OpenAI(base_url="http://127.0.0.1:<port>/v1", api_key="x")
# Pass-through to Ollama
r = c.chat.completions.create(
model="ollama/qwen3:14b-iq4xs",
messages=[{"role": "user", "content": "Hi"}],
)
# Orchestrated: a recipe (system prompt + tools + model), transparent to the
# client. The chat-loop happens inside woollama; client sees only the final answer.
r = c.chat.completions.create(
model="woollama/streamer",
messages=[{"role": "user", "content": "Please count to 4."}],
)
# A device-aware inferencer (declares `management_url`): "default" resolves to
# whatever's currently loaded, and woollama loads/queues on demand β no
# not-loaded error, no wedge.
r = c.chat.completions.create(
model="device/default",
messages=[{"role": "user", "content": "Hi"}],
)
# Pass-through also covers images and embeddings:
img = c.images.generate(model="device/Z-Turbo", prompt="a small blue teapot")
vec = c.embeddings.create(model="device/Embedder", input="hello world")
```
woollama serves on **two transports at once**: a Unix socket at
`$XDG_RUNTIME_DIR/woollama.sock` (mode 0600 β the default for local MCP clients,
since a connectable socket can spend the router's API keys) and an ephemeral
loopback TCP port written to `$XDG_RUNTIME_DIR/woollama.addr` for clients to
discover. The `<port>` above is that ephemeral port. Same pattern as a local
`fabric --serve` instance.
## Install
The router is **`woollamad`** β a small Rust daemon. The Python implementation is
kept as a reference server and the differential-test oracle (see below), but
`woollamad` is the canonical router.
**From crates.io** (once published β `cargo install` ships only the binary, so
bring your own `mcp.json`):
```sh
cargo install woollama-server # installs the `woollamad` binary
woollamad # starts the router; prints its address
```
**From this checkout** (works today; includes the bundled example MCP servers):
```sh
git clone https://github.com/teaguesterling/woollama
cd woollama
cargo build --release # builds target/release/woollamad
./target/release/woollamad # starts the router; prints its address
```
On startup `woollamad` prints its `OpenAI base_url` (e.g.
`http://127.0.0.1:<port>/v1`) β copy that into your OpenAI client. (It's also
written to `$XDG_RUNTIME_DIR/woollama.addr` for programmatic discovery, and it
serves the same surface over the `woollama.sock` unix socket.)
### Python 3.14: you will build from source, slowly
The `woollama` **Python** package depends on `woollama-core`, a compiled extension. We ship
wheels for **CPython 3.11β3.13 only** β pyo3 could not build against 3.14 until recently, and the
cp314 wheels are not published yet (#43).
Our metadata says `requires-python = ">=3.11"` with no upper bound, which is true β 3.14 works β
but on 3.14 there is no wheel, so the install falls back to **building the extension from
source**. That needs a Rust toolchain and takes minutes rather than seconds.
The trap is that `uv` picks an interpreter for you. `uv sync` in a project that merely allows
3.11+ will happily select 3.14 and then do a source build, or fail outright if no Rust toolchain
is present β and the error names `python-source` and maturin, neither of which mentions us. If
you want wheels today, pin the interpreter:
```sh
uv sync --python 3.12 # or any of 3.11-3.13
```
Any project depending on `woollama-core` **below 0.9.0** should pin 3.11β3.13 outright: those
versions cannot build on 3.14 at all, and every `woollama-core` sdist at or below 0.8.1 is
unbuildable on *any* Python (fixed in 0.8.2 β see the changelog for v0.14.4).
### The Python reference server
The original Python implementation still runs and is used as the live oracle that
keeps `woollamad` honest:
```sh
uv sync # creates .venv and installs deps
uv run woollama # the Python reference server
```
> **Prerequisite for the examples below:** they use `ollama/qwen3:14b-iq4xs`, so
> install [Ollama](https://ollama.ai), `ollama serve`, and
> `ollama pull qwen3:14b-iq4xs`. **No Ollama?** Use the keyless Claude path
> instead β `model="claude-code/haiku"` (needs the `claude` CLI logged in) β or
> any cloud model with its key set (see [Configuration](docs/configuration.md)).
### Tests & lint
```sh
# Rust (woollamad): the daemon's own suites
cargo test --tests --features test-fixtures
cargo build --release # so the live oracle can spawn the binary
# Python: hermetic suite + lint
uv run --extra dev pytest # hermetic suite (live tests are opt-in: -m integration)
uv run ruff check . # lint β the CI gate
# The live differential oracle β same tests, against woollamad by default:
uv run --extra dev pytest -m integration # targets target/release/woollamad
WOOLLAMA_TEST_CMD="python -m woollama" \
uv run --extra dev pytest -m integration # opt in to the Python reference
```
CI (`.github/workflows/ci.yml`) runs the Rust + Python gates on every push to `main` and PR.
For the same lint gate locally on commit, opt into the pre-commit hook:
```sh
uv tool install pre-commit && pre-commit install
```
Lint only β the project does not use `ruff format` (lines are hand-wrapped,
`E501` is ignored), so there is no formatter step in either gate.
## Design principles
1. **Two standards, neither extended.** MCP for tool/prompt/resource
discovery and execution; OpenAI chat-completions for the inference
primitive. woollama is a router between them.
2. **Local-only, ephemeral by default.** Random loopback port, persisted
address file for discovery, never `0.0.0.0` without explicit opt-in β and
the opt-in requires an auth token (`WOOLLAMA_TOKEN`; woollama refuses to
start off-loopback without one). The router holds API keys and routes to
local resources β it should not be LAN-reachable unauthenticated.
3. **The model namespace is the universal addressing scheme.** Raw inferencers
(`<provider>/<model>`, e.g. `ollama/X`, `anthropic/X`, `claude-code/X`) and
full recipes (`woollama/<recipe>`) are all addressable through OpenAI's
standard `model` field. No new wire format.
4. **woollama owns routing, not inference or tools.** It uses other people's
inference engines (Ollama, Anthropic, β¦) and other people's tool servers
(any MCP server β filesystem, git, lackpy, β¦). It composes them.
5. **she talks to llamas.**
## What works today
- OpenAI surface: `/v1/models`, `/v1/chat/completions` (pass-through +
recipe orchestration, both with `stream:true` β OpenAI SSE), plus
`/v1/images/generations` and `/v1/embeddings` pass-through (see below).
(`/v1/tools` introspection is Python-reference-only today β see
[#23](https://github.com/teaguesterling/woollama/issues/23))
- **Stateful surface**: `/v1/responses` (stateless subset, incl. `stream:true` β
OpenAI Responses SSE, + stateful) and `/v1/conversations` (create/list/get/
delete, plus `items` where the backend exposes its transcript). woollama routes
conversation *handles*; backends own state (woollama never stores transcripts
itself) β `claude-resume` for `claude-code` models, `managed-agents` (Anthropic
Managed Agents) for `claude-agent` models, with an interactive
`requires_action` pause/answer path; models with no state-owning backend are
stateless (`store:false`)
- Multi-backend routing by `<provider>/<model>`: ollama (incl. `num_ctx` honored
via ollama's native `/api/chat`), anthropic, openai, groq, together,
openrouter, `claude-code`, + any OpenAI-compatible endpoint via
`inferencers.toml`
- **Tool delegation**: a `claude-code` recipe with tools runs as an *executor* β
Claude owns the agentic loop and calls the recipe's allow-listed MCP tools
itself (per-recipe `--mcp-config` + `--allowedTools` containment)
- **Image + embedding pass-through**: `POST /v1/images/generations` and
`POST /v1/embeddings` forward a `<provider>/<model>` request to that
inferencer's own OpenAI-compatible endpoints, mirroring the chat pass-through
(non-streaming; unknown namespace β `400`)
- **Model pooling / device-aware inferencers**: an inferencer that declares a
`management_url` gets on-demand model loading (no more bare not-loaded
errors), stable **virtual model names** (`<provider>/default`, config
`virtual` aliases), and per-model request queuing with `503` + `Retry-After`
backpressure instead of a wedge β plus queue-aware LRU eviction to fit
`pool_max` loaded models. Fully additive: an inferencer with no
`management_url` is unaffected. Ships in **both** `woollama` (Python) and
`woollamad` (Rust); pooling currently covers `/v1/chat/completions` only
(`/v1/responses` isn't pooled yet)
- MCP server side: stdio (`woollamad mcp`) **and** Streamable HTTP at `/mcp` on
the same port β recipes as **parameterized prompts** (their `{{var}}` tokens β
arguments), a `chat` verb (with live tool-progress notifications), and every
downstream tool re-exported with its `output_schema` (aggregator)
- **MCP client side, stdio *and* HTTP**: an `mcp.json` server is either a
`command` (stdio subprocess) or a `url` (Streamable HTTP). Since woollama also
*serves* `/mcp`, the `url` form is how **one woollamad consumes another** β an
inference-holding instance reaching a tools-only instance's namespace, so the
tools instance can hold zero provider API keys. Credentials are `${VAR}` in
`headers`, validated fail-closed at load. `woollamad check-config` validates
every config file and exits non-zero on any problem
- **Pattern templating** on woollama's own `/w1/` namespace (not OpenAI's `/v1/`):
parameterized recipes/patterns with `{{var}}` substitution β
`GET /w1/patterns` (discovery), `POST /w1/patterns/{name}/render` (assemble),
`POST /w1/patterns/{name}/run` (render + infer, streaming). Patterns also come
from a fabric-style directory scan (`[patterns]`) and a **fabric backend**:
woollama can run/own `fabric --serve`, surface its library on `/w1/`, and
transparently proxy fabric's API at `/fabric/*`. Pattern backends are pluggable
(the `PatternBackend` trait β see `docs/extending.md`)
- File-driven config (`mcp.json`, `recipes.toml`, `inferencers.toml`), multi-
MCP-server discovery + unified tool registry, long-lived MCP connections
- Recipe allow-list enforced as a security boundary (in-loop AND in delegation);
served on a **Unix socket + loopback TCP**, address discovery file; CI
(ruff + hermetic suite, 3.11/3.12)
## Not yet (next on the roadmap)
- The live, interactive Claude-in-tmux session backend (a separate Rust session
driver) β gated on spikes that need a real terminal. (The interactive
`requires_action` path itself already works via the managed-agents backend.)
- cosmic-fabric actually consuming the conversations surface (the last open
integration milestone). The generic `store-backed` mechanism + two reference
store providers (MCP + REST) already ship; what's pending is the cross-repo
wiring. (Pattern templating + the fabric backend it needed have **shipped** β
see `docs/patterns.md`.)
- lackpy re-pinning to the now-published `woollama-core` wheel.
Full scorecard, ordering, and pending verifications:
**[`docs/roadmap.md`](docs/roadmap.md)**.
## Origin
woollama is the production-grade rewrite of an architecture co-designed
in [cosmic-fabric](https://github.com/teaguesterling/cosmic-fabric), which
remains a frontend (and will use woollama as its router engine). The design
docs that brought woollama here:
- `docs/architecture.md` β the model/tool/executor router design
- `docs/naming.md` β how we landed on this name
## License
MIT β see `LICENSE`.
This server cannot be deployed
Maintenance
ActivityActive
ResponsivenessResponsive