Taste-Context Brief
by weioai
README.md
# Taste-Context Brief (built on Qloo)
A tool that turns a small local business's **public site summary** plus a few **cultural seed terms** into a taste-context brief. In live mode Qloo supplies cross-domain taste entities (music, film and TV, brands, places) for those seeds; messaging ideas are written from them, **every one labelled as an inference**. In fixtures mode the entities are hand-written samples and are labelled as such everywhere.
It is an agentic tool: an MCP server that any MCP-capable agent can call, wrapping a fixed, deterministic tool pipeline. It is not itself an autonomous agent: the model never chooses a Qloo call and never reacts to Qloo results (in the default recorded mode no model runs at all).
It is one Cloudflare Worker with three interfaces:
| Interface | Where |
|---|---|
| HTML demo page | `GET /` |
| JSON API | `GET /api/scenarios`, `POST /api/brief`, `GET /healthz` |
| MCP server (Streamable HTTP) | `POST /mcp` |
Hosted demo (fixtures mode): https://audience-brief.weio.ai. Source: https://github.com/weioai/qloo-audience-brief. The examples below use `$BASE`: set it to the hosted demo or to your own deployment (for example the address `wrangler dev` prints).
Not affiliated with or endorsed by Qloo. The package, Worker and MCP tool ids (`qloo-audience-brief`, `weio-audience-brief`, `audience_brief`) are historical identifiers kept stable. The product name avoids leading with Qloo's name (it would suggest a Qloo product) and avoids the word "audience" (this tool makes no claims about anyone's audience).
## Status (read this first)
Live Qloo calls are implemented from Qloo's public documentation, but they have never been run against the real Qloo API: Weio does not hold an authorized Qloo hackathon API key yet. They are tested against a fake `fetch`, and the request path (including the no-redirect rule) was smoke-tested on a local `workerd` build against local stub servers, not against Qloo.
The hosted demo at https://audience-brief.weio.ai runs on labelled fixture data until Weio holds a key; live mode is not switched on there.
In fixtures mode every Qloo response is a hand-written sample in Qloo's documented response format. All named entities are fictional (for example "Example Ramen Bar (fixture)"), every entity id contains `fixture`, and every response says `data_source: "fixture"` and carries the fixture note, both at the top level and inside `brief.qloo_facts`. Fixture output never says that Qloo supplied anything. Nothing is invented for custom input: the fixture layer answers only for the three built-in scenarios and otherwise answers `409`.
## The problem, the pipeline, the output
**Problem.** A small business wants website copy that feels culturally "of a piece" with its neighbourhood and taste, but has no customer data to mine and should not be guessing about people.
**Pipeline** (`src/agent.mjs`, with a step-by-step trace in every response). A fixed sequence of tool calls, written in code:
1. **Plan.** Use the given seed terms (or the scenario's) and choose the target domains: music (`urn:entity:artist`), film (`urn:entity:movie`), brands (`urn:entity:brand`) and places (`urn:entity:place`, only when a city is given). With no seeds, and only when a model is enabled, the model proposes 1-3 word terms from the summary and the request ends there: they come back in a `422` as `proposed_seeds` for you to review and resubmit. Terms derived from your summary are never sent to Qloo unreviewed.
2. **Resolve** each seed term. Topics (genres, foods, hobbies) go to `GET /v2/tags?filter.query=<term>` and resolve to a tag id. If no tag matches, or the tag lookup is refused, a named thing (an artist, a film, a brand) is looked up with `GET /search` (top match; Qloo's documented `404 No results` counts as no match). A seed costs one upstream call, or two when the entity search is needed.
3. **Insights** per domain with `GET /v2/insights`, using resolved tags as `signal.interests.tags` and resolved entities as `signal.interests.entities` (the city goes only to the place query, as `filter.location.query`). Seeds that resolve to the same id are merged.
4. **Render the brief** (`src/brief.mjs`): bounded, with Qloo facts kept separate from everything the model writes, and carrying its own `data_source` label.
5. **Narrative.** Up to four messaging ideas, each citing the entities that inspired it and ending in `(inference)`. The text comes from a replay (recorded mode), a deterministic template, or a model. For model output the Worker checks every `Inspired by:` list against the entities actually in the brief and drops any idea that cites something else; if no idea survives it falls back to the template and says so in the note.
6. **Return** `{mode, data_source, brief, narrative, trace, usage, disclosure}`. The `disclosure` is built from what produced the data and the text in that response (fixture data, template text, a replay, or a named model).
**Output.** A brief with the entities per domain (affinity where Qloo gave one), a narrative labelled `inference`, and the trace of what the pipeline did. The entities are cultural context; the brief says in so many words that they are *not evidence about this business's customers*.
A brief spends at most **6 upstream calls** (enforced by a counter). The planner trims domains so up to three seeds can be resolved (with three or more seeds, a city adds the place domain and drops film). A seed that does not fit in the budget, including one that needs the second lookup, is listed under `qloo_facts.seeds_not_resolved`. Each upstream call has a 10 s timeout and never follows redirects. A failed Qloo call is reported as a failure; it is never replaced with made-up data.
**Not built:** an "LLM-only answer versus Qloo-grounded brief" comparison; audience or demographic signals (deliberately left out: this tool makes no claims about customers); a daily spend ceiling (see Spend control).
## Using it
### Web page
Open the page, pick a sample business, and read the facts table, the narrative (with its **Inference** badge) and the tool-call trace. The custom-input form works in live mode; in fixtures mode it shows the server's honest `409` message.
### JSON API
```sh
BASE=https://audience-brief.weio.ai # hosted demo (fixtures mode), or your own deployment's URL
# built-in sample inputs
curl -s $BASE/api/scenarios
# run a built-in scenario
curl -s $BASE/api/brief \
-H 'content-type: application/json' \
-d '{"scenario":"bookstore-cafe"}'
# custom input (live mode only; fixtures mode answers 409)
curl -s $BASE/api/brief \
-H 'content-type: application/json' \
-d '{"site_summary":"A small tea house with live folk nights.","seeds":["folk music","tea"],"city":"Portland"}'
```
Request fields (all optional, but you need either `scenario` or `site_summary` plus `seeds`; unknown fields are rejected):
| Field | Rule |
|---|---|
| `scenario` | one of `ramen-shop`, `bookstore-cafe`, `yoga-studio` |
| `site_summary` | text, 1-800 characters; never sent to Qloo |
| `seeds` | 1-5 terms, 1-100 characters each; emails, links and phone numbers are rejected; sent to Qloo as typed |
| `city` | text, up to 80 characters, no digits, emails or links; used only for place results |
Status codes: `200` ok, `400` invalid input, `409` custom input in fixtures mode, `413` body over 8 KB, `415` not `application/json`, `422` no seed term matched in Qloo, or (no seeds given, model enabled) the response carries `proposed_seeds` and nothing was sent to Qloo, `429` rate limit exceeded (with `Retry-After`), `502` Qloo request failed (generic message plus upstream status; upstream text is never echoed), `503` live mode without a configured key, or a spending mode without a rate limiter.
### MCP
Endpoint: `$BASE/mcp` (Streamable HTTP, JSON responses, stateless; `GET /mcp` answers `405`). Protocol versions 2025-06-18, 2025-03-26 and 2024-11-05 are accepted; a request whose `MCP-Protocol-Version` header names any other version gets `400`, and a request without the header is treated as 2025-03-26. JSON-RPC batches of up to 4 messages are answered in order.
Claude Code:
```sh
claude mcp add --transport http audience-brief $BASE/mcp
```
Generic client configuration for clients that support remote HTTP servers (this is the hosted demo; use your own host if you deploy it yourself):
```json
{
"mcpServers": {
"audience-brief": {
"type": "http",
"url": "https://audience-brief.weio.ai/mcp"
}
}
}
```
Clients that only speak stdio can bridge with a remote-MCP proxy such as `mcp-remote`:
```json
{
"mcpServers": {
"audience-brief": {
"command": "npx",
"args": ["mcp-remote", "https://audience-brief.weio.ai/mcp"]
}
}
}
```
Raw JSON-RPC works too:
```sh
curl -s $BASE/mcp \
-H 'content-type: application/json' \
-H 'accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"audience_brief","arguments":{"scenario":"ramen-shop"}}}'
```
Tools (both annotated `readOnlyHint: true, openWorldHint: true`):
- `list_scenarios` (no arguments): the built-in sample businesses, plus the current mode and, in fixtures mode, the fixture note.
- `audience_brief` (`scenario?`, `site_summary?`, `seeds?`, `city?`): the same brief as `POST /api/brief`. The result is human-readable text that includes the data-source label, plus `structuredContent` with the full JSON. Tool failures (custom input in fixtures mode, live mode not configured, no seed match, proposed seeds to confirm, rate limited) come back as `isError: true`; malformed arguments are JSON-RPC `-32602`. The model-written ideas in the text are marked as untrusted model output.
## Modes
| Variable | Value | Meaning |
|---|---|---|
| `MODE` | `fixtures` (default) | Qloo responses come from hand-written, labelled fixtures; only the 3 built-in scenarios work |
| | `live` | Real calls to the Qloo hackathon API with `QLOO_API_KEY`. Without the key every brief request is `503 {"error":"live mode is not configured"}`; it never falls back to fixtures. Needs a rate limiter (see Spend control) |
| `LLM_MODE` | `recorded` | Replays the narrative stored with a built-in scenario (fixtures mode only; in live mode it becomes `template`) |
| | `workers-ai` | Calls the model through the Workers AI binding, only after the model passes the allow-list **and** `WORKERS_AI_APPROVED = "yes"` is set (see Models); on any failure the narrative falls back to the template and says so |
| | `template` (default for unknown values) | Deterministic text built from the brief; no model involved |
Unknown values fall back to `fixtures` / `template`. `narrative.source` always says which of `recorded`, `workers-ai` or `template` produced the text, and `data_source` is `fixture` or `qloo-live`.
Other variables: `QLOO_API_KEY` (secret), `LLM_MODEL` (default `@cf/meta/llama-3.3-70b-instruct-fp8-fast`), `REPO_URL` (footer link, https only; with no value the page says the repository is not published yet), `WORKERS_AI_APPROVED`. Bindings: `RATE_LIMITER`, `GLOBAL_RATE_LIMITER` (optional), `AI` (optional, see Models).
## Privacy
Sent to **Qloo**: the seed terms exactly as given (as `/v2/tags` and `/search` queries), and for place results an optional city name, plus the API key as an `X-Api-Key` header. Seed terms that look like an email, link or phone number are rejected before any request, and the city is held to the same screen and may not contain digits. These checks cannot recognise a personal or business name: **a name typed as a seed is sent as typed, so keep names out of seeds.**
**Never** sent to Qloo: the site summary, the business name, visitor identifiers, IP addresses, and anything a model derives from your summary. Seeds proposed by a model are returned to you for review; they are not sent to Qloo.
Sent to the **model** (only in `workers-ai` mode): the public site summary and the names and affinity scores of the entities. Angle brackets in that text are neutralised so it cannot close the prompt's `<site_summary>` fence. In `recorded` and `template` modes nothing leaves the Worker. The recorded narratives were generated once, offline, from the fictional fixture data, not from any real business.
The Worker keeps no storage, sets no cookies, and does not log request bodies. Requests are capped at 8 KB and answered with a strict Content-Security-Policy (`default-src 'none'`, scripts only from `'self'`, no inline script); the page inserts all dynamic text with `textContent`.
## Spend control
Any request that can cost Qloo quota or model usage (`MODE = "live"`, or `LLM_MODE = "workers-ai"`) is refused with `503` unless a `RATE_LIMITER` binding (a Workers Rate Limiting binding) exists, so a public deployment cannot run unlimited. Each client (keyed on `CF-Connecting-IP`) has its own budget and is refused with `429` when it is spent; an optional `GLOBAL_RATE_LIMITER` binding caps total traffic. A limiter that errors or answers oddly also refuses the request. Fixtures mode with the recorded or template narrative costs nothing and needs no limiter. Each tool call inside an MCP batch is limited on its own.
What this is not: the limits count per Cloudflare location over a 10 or 60 second period, so they are rate limits, not a daily total; there is no daily ceiling. A model call that hits the 25 s timeout is not cancelled and can still finish. CORS is `*` on `/api/*` and `/mcp`; it does not restrict non-browser callers, so the limiter, not CORS, is the spend control.
## Models
US-provider models only. The default is **Meta Llama 3.3 70B** (`@cf/meta/llama-3.3-70b-instruct-fp8-fast`) through Cloudflare Workers AI. Every model id passes `checkModel()` in `src/llm.mjs` before any call: a deny-list of foreign model families is checked first, then the vendor prefix must belong to an allowed family (Anthropic, OpenAI, Google, Meta, NVIDIA, Microsoft, Amazon). The system prompt forbids claims about customers, visitors or demographics, requires each idea to cite the supplied entities by name and be labelled an inference, and limits the answer to 180 words of plain text.
Workers AI is called from inside the Worker through the `env.AI` binding. If your organisation's model policy requires such calls to go through a specific wrapper or approval, do not enable this route. `LLM_MODE = "workers-ai"` is ignored (the Worker uses the template and `/healthz` says why) unless `WORKERS_AI_APPROVED = "yes"` is also set, to record that your policy allows this route. It also needs the rate limiter above.
The recorded narratives were produced once from a real model run (Meta Llama 3.3 70B through Workers AI) over the fixture data, with the current prompt, and are replayed exactly as written. A test checks that every idea in them cites only entities that are in its own fixture brief, the same check applied to live model output.
## Built by AI
Built by AI: this project was designed and written by AI agents (Anthropic Claude and OpenAI Codex) working for Weio, Inc.; a human owner is accountable for it. Not affiliated with or endorsed by Qloo.
The original Python scaffold (a dependency-free Qloo adapter and an offline brief renderer) was written by an OpenAI Codex agent. An Anthropic Claude agent ported it to JavaScript and extended it into the pipeline, the Worker, the MCP server and the demo page. The Qloo request shapes follow Qloo's public developer documentation, with the places where those docs are inconsistent or thin isolated and commented in `src/qloo.mjs`.
Build dates: all of the code, tests and documentation here were created in October 2026, starting on 2026-10-02; none of it existed before.
## Develop
Requires Node 18 or newer. There are no npm dependencies and no build step; `src/` uses Web APIs only.
```sh
npm test # node --test test/*.test.mjs
```
The tests call the Worker's `fetch` directly and cover the routes, each scenario end to end, fixtures versus live behaviour, the request shapes sent to Qloo (via a recording fake `fetch` that, like the Workers runtime, rejects any `redirect` value other than `follow` or `manual`), the model allow-list, citation checking, the MCP protocol, request limits, rate limiting and security headers. No test touches the network.
```
src/worker.mjs Workers entry: exports only the default handler
src/app.mjs routes, limits, rate limiting, headers, CORS
src/config.mjs env -> {mode, llmMode, ...}
src/agent.mjs the fixed pipeline and its trace
src/qloo.mjs Qloo client + live transport (call cap, timeout, no redirects)
src/fixtures.mjs fixture transport (same interface as live)
src/fixtures-data.mjs hand-written fictional data + recorded narratives
src/brief.mjs validation and brief rendering
src/llm.mjs prompt, allow-list, citation check, recorded / workers-ai / template
src/mcp.mjs MCP over Streamable HTTP
src/page.mjs HTML page and /app.js
```
## Deploy
```sh
npm install --global wrangler # or: npx wrangler
wrangler deploy # fixtures mode, no secrets needed
```
`wrangler.toml` ships with `MODE = "fixtures"` and `LLM_MODE = "recorded"`. After deploying, check that the demo URL and the repository link both resolve before pointing anyone at them.
To enable live mode (only with an authorized Qloo hackathon key): store the key with `wrangler secret put QLOO_API_KEY`, add a `RATE_LIMITER` rate-limiting binding as shown in the comments in `wrangler.toml`, set `MODE = "live"`, and set `LLM_MODE = "template"`. Without the binding every live request is refused. Do not turn on `workers-ai` until your model policy allows calling Workers AI from inside a Worker (see Models).
## License
MIT, see [LICENSE](LICENSE). Copyright (c) 2026 Weio, Inc. Not affiliated with Qloo.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues