Salla MCP Server
by jamalla
README.md
# Salla MCP Demo: contract-driven Code Mode
A **teaching project** showing how to build a scalable MCP server for a real e-commerce API and an agent that uses it. It has four parts:
| Part | Folder | Command | What it does |
|---|---|---|---|
| **MCP server** | `src/mcp_server/` | `uv run salla-mcp serve` | One generic engine for the [Salla](https://salla.sa) Admin API. Everything it can do comes from **task contracts**: YAML and Markdown files that partners add **without writing code**, including the screens the agent can show (**A2UI views**). |
| **Chat agent** | `src/chat_agent/` | `uv run salla-agent` | An OpenAI-powered web chat that uses the server. You approve every change it wants to make to your store. |
| **Task Contract Studio** | `src/task_studio/` | `uv run salla-studio` | A form for building, previewing, validating and installing task contracts on the running server. |
| **A2UI renderer** | `src/a2ui_renderer/` | (served by the two web apps) | A small plain-JavaScript renderer for [A2UI v0.9.1](https://a2ui.org) screens. |
> ⚠️ This project calls the **real** Salla API. Approved changes modify a real store, so use a **Salla demo store** while learning.
**Contents:**
1. [The idea in one picture](#1-the-idea-in-one-picture)
2. [Reproduce it exactly](#2-reproduce-it-exactly)
3. [Configuration reference](#3-configuration-reference)
4. [How it works](#4-how-it-works)
5. [Built for 100+ tasks](#5-built-for-100-tasks)
6. [Project tour](#6-project-tour)
7. [Build a similar project from scratch](#7-build-a-similar-project-from-scratch)
8. [Tests](#8-tests)
9. [Troubleshooting](#9-troubleshooting)
10. [Versions, limits and history](#10-versions-limits-and-history)
Deeper reading: [docs/architecture.md](docs/architecture.md) (flows, safety, scaling) and [docs/adding-a-task.md](docs/adding-a-task.md) (the partner guide).
---
## 1. The idea in one picture
```mermaid
flowchart LR
subgraph P["Partner files (data, no code)"]
T["tasks/categories/<br/>task.yaml · functions.yaml · ui.yaml<br/>prompts/*.md · skills/*/SKILL.md"]
end
subgraph E["mcp_server engine (code, written once)"]
C["catalog<br/>(validated)"] --> S["search · describe"]
C --> X["call · execute<br/>(Monty sandbox)"]
X -->|reads| H["HTTP executor"]
X -->|writes| PL["change plan"]
PL --> A["apply_changes<br/>ONE approval gate"] --> H
C --> V["render_ui<br/>(A2UI views)"]
end
T --> C
H <-->|HTTPS| API["Salla Admin API"]
AG["chat_agent<br/>(OpenAI)"] <-->|"MCP · Streamable HTTP"| E
B["Browser<br/>(draws A2UI)"] <-->|WebSocket| AG
ST["task_studio"] -->|"admin API"| C
```
Five ideas carry the whole design:
1. **Tasks are data, the engine is code.** Adding a business domain means adding a folder of YAML and Markdown. The server, the agent and the Studio don't change.
2. **Code Mode.** The LLM gets 6 fixed tools, whatever the number of tasks. It finds functions with `search`, reads their typed signatures with `describe`, and combines them in short Python scripts that run in a secure sandbox (Monty).
3. **One approval gate, one HITL broker.** Every write is *staged* into a change plan. Nothing changes in the store until the human approves the whole plan through MCP elicitation. On the agent side, one small file, the [HITL broker](#44-human-in-the-loop-one-gate-in-the-server-one-broker-in-the-agent), handles every interrupt and resume, for any task.
4. **Planned screens.** Partners design each screen ahead of time in A2UI. The LLM only fills in the data, and button clicks come back to the LLM.
5. **Built for scale.** The prompt sent on every LLM call doesn't grow with the tasks; everything else is loaded on demand. A CI test with 100 generated tasks enforces this.
---
## 2. Reproduce it exactly
### 2.1 The environment it was built and tested with
| Item | Version used | Notes |
|---|---|---|
| OS | Windows 11 Pro | Commands below are given for bash and PowerShell |
| Python | 3.13.12 | `pyproject.toml` requires ≥ 3.11. uv downloads a suitable Python if you have none |
| [uv](https://docs.astral.sh/uv/) | 0.11.29 | Installs the exact dependency versions from `uv.lock` |
| Node.js | 22.14 | *Optional*, only to run an MCP inspector with `npx` |
| Browser | Edge / Chrome | Needs internet access for two CDN scripts: marked 12.0.2 and DOMPurify 3.1.6 |
Key libraries, pinned in `uv.lock`:
| Library | Version | Used for |
|---|---|---|
| `mcp` | 2.2.0 | MCP Python SDK: `MCPServer` (server) and `Client` (agent), Streamable HTTP, elicitation |
| `pydantic-monty` | 0.0.23 | The Code Mode sandbox (pre-1.0, so pinned exactly) |
| `openai` | 3.13.0 | Responses API (the agent's LLM) |
| `fastapi` / `starlette` / `uvicorn` | 0.141.1 / 1.6.0 / 0.52.4 | The chat and Studio web apps, WebSocket |
| `httpx` | 0.28.1 | Calls to the Salla API |
| `pydantic` / `pydantic-settings` | 2.13.5 / 2.15.0 | Contract models, settings from `.env` |
| `pyyaml` | 6.0.3 | Contract files (uses libyaml's C parser when available) |
| `jsonschema` / `referencing` | 4.26.0 / 0.37.0 | Argument validation, and A2UI schema checks |
| `pytest` / `pytest-asyncio` / `respx` | 9.1.1 / 1.4.0 / 0.23.1 | Tests, with the Salla API mocked |
Protocols and services:
- **MCP** over Streamable HTTP;
- **A2UI v0.9.1** with the basic catalog (the official schemas are copied into `src/mcp_server/engine/a2ui_spec/`);
- **OpenAI Responses API** with the model `gpt-5.5`;
- **Salla Admin API v2** (`https://api.salla.dev/admin/v2`).
### 2.2 Accounts you need
1. **A Salla demo store and an access token.**
1. Create a free account on the [Salla Partners portal](https://portal.salla.partners), and create a **demo store**.
2. Create an **app** with the scopes **`categories.read`** and **`categories.read_write`**.
3. Install the app on the demo store and get its **access token**. See Salla's [authorization guide](https://docs.salla.dev/doc-421118). In *Easy mode*, the token arrives in the `app.store.authorize` webhook event.
4. Tokens expire after about 14 days. A `401` means it's time for a new one.
2. **An OpenAI API key** with access to `gpt-5.5`, the default. Other models that support the Responses API and function calling can be set with `OPENAI_MODEL`, but only `gpt-5.5` was tested.
The tests need **neither**: Salla is mocked, and the LLM is scripted.
### 2.3 Step by step
**1. Get the code and install the exact versions.**
```bash
git clone https://github.com/jamalla/salla-api-operation-mcp-task.git
cd salla-api-operation-mcp-task
uv sync --locked # creates .venv with the versions from uv.lock
```
**2. Create your `.env`**, from the repository root. Every app reads `./.env` from the folder it's started in.
```bash
cp .env.example .env # PowerShell: Copy-Item .env.example .env
uv run python -c "import secrets; print(secrets.token_hex(24))" # a value for MCP_ADMIN_TOKEN
```
Fill in at least these three:
```ini
SALLA_ACCESS_TOKEN=... # from 2.2
OPENAI_API_KEY=sk-... # from 2.2
MCP_ADMIN_TOKEN=... # the random value above; the Studio needs it
```
**3. Check the contracts and run the tests.** Neither needs a token.
```bash
uv run salla-mcp validate # → OK categories: 8 functions, 3 prompts, 1 skills, 2 views
uv run pytest # → 90 passed
```
**4. Start the three apps**, each in its own terminal, from the repository root:
```bash
uv run salla-mcp serve # MCP server → http://127.0.0.1:8001/mcp (+ /admin/tasks)
uv run salla-agent # chat agent → http://127.0.0.1:8000
uv run salla-studio # the Studio → http://127.0.0.1:8020 (optional)
```
The MCP server prints:
```
Salla MCP server → http://127.0.0.1:8001/mcp
Profiles → none (set MCP_PROFILES)
Admin API → http://127.0.0.1:8001/admin/tasks
```
**5. Open http://127.0.0.1:8000** and try:
- **"List my hidden categories"**: `search` → `describe` → `execute`, and no approval is needed.
- **"Create Pets with a sub-category Cat food"**: one approval card with 2 changes. Try **Reject**, then **Approve**.
- **"Show my hidden categories as cards"**: the `category_list` view is drawn.
- Click **Details**: the detail view appears.
- Edit the name and click **Save name**: an approval card shows the new name.
- In the sidebar, **Prompts → Build a category tree**.
- Write in Arabic: «اعرض الأقسام المخفية».
Under each answer, the chat shows what the turn cost, for example: *4 LLM calls · 7,493 input tokens (1,792 cached) · 252 output*.
### 2.4 Check that it works
| Check | Expected |
|---|---|
| `uv run salla-mcp validate` | `OK categories: 8 functions, 3 prompts, 1 skills, 2 views` |
| `uv run pytest` | `90 passed` |
| `curl -H "Authorization: Bearer <MCP_ADMIN_TOKEN>" http://127.0.0.1:8001/admin/tasks` | JSON with `"ok": true` and the `categories` task, including its 8 functions and 2 views |
| Chat sidebar | `connected · gpt-5.5`; the 6 MCP tools; the `categories` functions and 2 `view` rows; 3 prompts; 1 skill |
| Studio (http://127.0.0.1:8020) | The brands template form, with a live **UI views** preview (two cards with a Rename button) |
| MCP inspector (optional) | `npx @mcpjam/inspector@latest`, or `npx @modelcontextprotocol/inspector`. Choose *Streamable HTTP*, URL `http://127.0.0.1:8001/mcp`, and you get 6 tools, 3 prompts and 2 resources |
---
## 3. Configuration reference
Every setting is an environment variable, usually set in `.env` (see [.env.example](.env.example)). The three apps read the same `.env`, from the folder they're started in.
**MCP server** (`src/mcp_server/config.py`):
| Variable | Default | Meaning |
|---|---|---|
| `SALLA_ACCESS_TOKEN` | *(required to serve)* | Token of the Salla app installed on your store |
| `SALLA_BASE_URL` | `https://api.salla.dev/admin/v2` | Salla Admin API |
| `TASKS_DIR` | `<repo>/tasks` | Where the task folders are |
| `ENABLED_TASKS` | *(empty = all)* | Comma-separated tasks to load, e.g. `categories` |
| `MCP_HOST` / `MCP_PORT` | `127.0.0.1` / `8001` | Streamable HTTP endpoint: `http://MCP_HOST:MCP_PORT/mcp` |
| `MCP_PAGE_SIZE` | `50` | Items per page for MCP lists and `search()` |
| `MCP_PROFILES` | `{}` | JSON, e.g. `{"catalog": ["categories"]}`. A client at `/mcp?profile=catalog` sees only those tasks |
| `MCP_ADMIN_TOKEN` | *(empty = admin API off)* | Bearer token for `/admin/tasks`, used by the Studio |
| `SANDBOX_MAX_SECONDS` | `10` | Time limit per `execute()` script |
| `SANDBOX_MAX_MEMORY_MB` | `128` | Memory limit per script |
| `PLAN_TTL_SECONDS` | `900` | Staged change plans expire after 15 minutes |
**Chat agent** (`src/chat_agent/config.py`):
| Variable | Default | Meaning |
|---|---|---|
| `OPENAI_API_KEY` | *(required)* | OpenAI key |
| `OPENAI_MODEL` | `gpt-5.5` | Model for the Responses API |
| `MAX_AGENT_STEPS` | `20` | Maximum LLM ⇄ tool round trips per user message |
| `MAX_TOOL_OUTPUT_CHARS` | `20000` | Longer tool results are cut before the LLM reads them |
| `MCP_SERVER_URL` | `http://127.0.0.1:8001/mcp` | The MCP server. Add `?profile=<name>` to use a profile |
| `AGENT_HOST` / `AGENT_PORT` | `127.0.0.1` / `8000` | The chat web app |
**Task Contract Studio** (`src/task_studio/app.py`):
| Variable | Default | Meaning |
|---|---|---|
| `MCP_ADMIN_TOKEN` | *(required to install)* | Must equal the server's value |
| `MCP_SERVER_URL` | `http://127.0.0.1:8001/mcp` | The admin API is derived from it: same host and port, path `/admin` |
| `MCP_ADMIN_URL` | *(derived)* | Set it only if the admin API is somewhere else |
| `STUDIO_HOST` / `STUDIO_PORT` | `127.0.0.1` / `8020` | The Studio web app |
**CLI** (`uv run salla-mcp …`):
| Command | Does |
|---|---|
| `serve [--host H] [--port P]` | Run the server (the default command) |
| `validate [task folders…]` | Check task folders against the contract; prints the file and field of every problem |
| `schema` | Regenerate `tasks/_template/*.schema.json` from the contract models |
---
## 4. How it works
### 4.1 What a partner writes (no Python)
A task is a folder: `tasks/<task>/` with `task.yaml`, `functions.yaml`, `ui.yaml` (optional), `prompts/*.md` and `skills/*/SKILL.md`. A function is a few lines of YAML:
```yaml
- name: create # exposed as categories_create
summary: Create a category, optionally under a parent.
effect: write # read | write | destructive: decides the approval
method: POST
path: /categories
params: # JSON Schema: validated, and turned into a typed signature
type: object
required: [name]
properties:
name: { type: string }
parent_id: { type: integer }
status: { type: string, enum: [active, hidden], default: hidden }
request:
body: { name: name, parent_id: parent_id, status: status }
approval: "Create category “{name}” ({status})"
```
`tasks/_template/` is the contract: JSON Schemas generated from the engine, plus commented examples. The server refuses to start on an invalid contract, and `uv run salla-mcp validate` shows the same messages. The full guide is [docs/adding-a-task.md](docs/adding-a-task.md).
### 4.2 Planned screens: `ui.yaml` (A2UI)
A task can plan the screens the agent shows, in [A2UI v0.9.1](https://a2ui.org). The partner writes each layout once; the LLM only sends the data:
```yaml
views:
- name: category_list # render_ui(view="categories_category_list", data=...)
summary: Show several categories as cards.
data: {type: object, required: [heading, categories], properties: {...}} # JSON Schema, validated
actions: # button events: a click goes back to the LLM
- {name: hide_category, label: Hide this category, function: categories_update}
components: # A2UI components, used exactly as written
- {id: root, component: Column, children: [heading, cards]}
- {id: cards, component: List, children: {path: /categories, componentId: card}}
# …
```
- **Checked at load.** Views must be valid A2UI, use only the components the bundled renderer can draw, and account for every binding and button.
- **Rendering.** `render_ui(view, data)` checks the data and returns three A2UI messages (`createSurface`, `updateComponents`, `updateDataModel`), which the chat draws.
- **Clicks.** A click sends an A2UI `action` back to the LLM as a `[UI action]` message. Any change that follows still needs your approval.
### 4.3 What the LLM sees: 6 tools, whatever the number of tasks
| Tool | Purpose |
|---|---|
| `search(query, kind, task, cursor)` | Ranked search (BM25) over functions, views, prompts and skills (paginated) |
| `describe(names)` | Exact Python signatures of functions, and the data schema of views |
| `call(function, arguments)` | Run **one** function deterministically |
| `execute(code)` | Run a short **Python script** in the Monty sandbox that combines functions (**Code Mode**) |
| `render_ui(view, data)` | Show a **planned screen** filled with data; also callable inside `execute` |
| `apply_changes(plan_id)` | **The only way the store changes.** The user approves or rejects the whole plan |
```python
# what an LLM sends to execute(): reads run now; writes are only staged
pets = await categories_create(name="حيوانات أليفة", en_name="Pets")
await categories_create(name="طعام قطط", en_name="Cat food", parent_id=pets["data"]["id"]) # "$w1.id"
```
This produces **one** approval card for 2 changes. On approve, the engine applies them in order, and `"$w1.id"` becomes the parent's real ID. The agent also has its own tool, `load_skill`, which reads a skill's full text on demand.
### 4.4 Human in the loop: one gate in the server, one broker in the agent
Every change to the store stops and waits for a human. This happens in two places, each written once and shared by every task:
| Where | File | Role |
|---|---|---|
| **The gate** (server) | [`api/approval.py`](src/mcp_server/api/approval.py), plus `apply_changes` in [`api/tools.py`](src/mcp_server/api/tools.py) | Decides *what* to ask. It renders the change plan as a table and asks through **MCP elicitation**. `apply_changes` has an `approval` parameter the LLM can't fill in; a resolver fills it with the human's answer. |
| **The HITL broker** (agent) | [`chat_agent/hitl.py`](src/chat_agent/hitl.py), about 70 lines | Handles *every* interrupt and resume. It pauses the tool call, asks the browser, and resumes the call with the answer. |
| The card (browser) | `approvalCard` in [`static/app.js`](src/chat_agent/static/app.js) | Draws any elicitation form from its JSON Schema. An `approve` field becomes Approve / Reject buttons; other fields become inputs. |
**The broker knows nothing about tasks.** It never sees categories, plans or Salla. It receives an elicitation (a message and the JSON Schema of a small form) and returns the human's answer: `accept` with the form data, `decline` or `cancel`. A new task, or a new kind of question from the server, goes through it with no change to the agent.
How one interrupt and resume works:
```mermaid
sequenceDiagram
participant L as Agent loop (agent.py)
participant C as MCP client
participant B as HITL broker (hitl.py)
participant W as Browser (app.js)
participant M as MCP server
L->>C: call_tool("apply_changes", plan_id)
C->>M: tools/call
Note over M: resolver confirm_plan returns Elicit(render_plan(plan), Approval)
M-->>C: elicitation request (message + form schema)
C->>B: elicitation_callback(params)
B->>W: approval_request (id, message, schema)
Note over L,B: the tool call is paused on an asyncio.Future
W->>B: approval_response (id, action, content), via main.py and broker.resolve
B-->>C: ElicitResult(accept, approve=true, note)
C-->>M: the answer
Note over M: approved: apply the plan in order. Otherwise nothing is sent to Salla
M-->>C: tool result
C-->>L: the loop resumes with the result
```
What makes it reliable:
- **One pending question, one future.** `elicitation_callback` creates a request id and an `asyncio.Future`, sends `approval_request` over the tab's WebSocket and awaits the future. `resolve(id, action, content)` completes it when the browser answers.
- **The chat keeps listening while it waits.** The agent turn runs as a background task (`main.py`), so the WebSocket loop can deliver the answer to the paused call.
- **Answers reach the right tab.** Each browser tab has its own MCP session and its own broker.
- **One question at a time.** The agent asks the LLM for one tool call at a time (`parallel_tool_calls=False`).
- **Nothing hangs.**
- A question expires after 10 minutes: the browser gets `approval_expired`, and the server gets `cancel`.
- Closing the tab cancels every open question (`cancel_all`).
- An unknown answer counts as `cancel`.
- URL-mode elicitation (e.g. "log in at this URL") is declined.
- **Fail closed.** Only an accepted form with `approve: true` applies the plan. A rejection, decline, cancel or timeout sends nothing to Salla, and the LLM is told why, so it doesn't retry. Plans are single-use and expire after `PLAN_TTL_SECONDS`.
**Reusing it.** The broker depends only on the MCP SDK's elicitation types and a `send(event)` function. To use it in another client, such as a CLI, a Slack bot or a mobile app, keep `hitl.py` and replace the transport: whatever shows the question calls `broker.resolve(...)` with the answer.
### 4.5 The Studio and the admin API
The Studio edits a form and renders it into contract files, shown as live YAML tabs with an A2UI preview for views. It sends them to the MCP server's **admin API**, which validates them with the start-up rules and installs them **live**, with no restart. The admin API runs next to `/mcp`, is protected by `MCP_ADMIN_TOKEN`, and is *not* part of MCP, so LLMs never see it.
| Endpoint | Purpose |
|---|---|
| `GET /admin/tasks` | List installed tasks |
| `GET /admin/tasks/{name}` | Get a task's contract files |
| `POST /admin/tasks/validate` | Dry run: `{"files": {"task.yaml": "...", ...}}` |
| `POST /admin/tasks` | Install and hot-reload (`"replace": true` to update) |
| `DELETE /admin/tasks/{name}` | Uninstall |
### 4.6 Pagination at three levels
1. **MCP lists** (prompts, resources) come in pages with an opaque `next_cursor` (`MCP_PAGE_SIZE`).
2. **`search()`** returns a `next_cursor`.
3. **Salla's own pages:** `fetch_all_pages("categories_list", {...})` inside `execute` walks them all.
---
## 5. Built for 100+ tasks
Every LLM call sends the agent's instructions and the tool schemas again, and one user message often makes 4–6 calls. So the rule is: **the fixed part of the prompt must not grow with the tasks; everything else is loaded when needed.**
| Sent on every LLM call | Loaded only when needed |
|---|---|
| 6 tools + `load_skill`, server and agent instructions, the task **names** | functions and views (`search` → `describe`), a skill's text (`load_skill`), prompts (when the user clicks one), data (tool results) |
Measured with 100 generated tasks (800 functions, 200 views, 100 skills):
| | Result |
|---|---|
| Fixed prompt, 1 task → 100 tasks | about 1,800 → 2,050 tokens, roughly 3 tokens per task |
| Listing every skill's description instead | about 5,700 extra tokens on every call |
| Search over 1,400 entries | about 3 ms |
| Server start-up with 100 tasks / reload after changing one task | about 7 s / 0.3 s |
How:
- **No skill list in the prompt.** The agent's prompt names the tasks; skills are found with `search(kind="skill")`.
- **Size limits in the contract.** Summaries are limited to 200 characters, action labels to 150, skill descriptions to 300, and a skill's text to 12,000.
- **Ranked search (BM25).** Rare words and words in names count more, so the LLM needs fewer search rounds.
- **Small results.**
- Scripts return only the fields they need.
- `render_ui` inside `execute` sends list data straight to the screen.
- The agent cuts results longer than `MAX_TOOL_OUTPUT_CHARS`.
- **Only the stubs a script uses.** `execute` type-checks a script against the functions it names, not against all of them.
- **Profiles.** `MCP_PROFILES` plus `…/mcp?profile=<name>` gives each kind of agent only its tasks, and this is enforced for every tool, prompt and resource.
- **Incremental reloads.** Installing a task re-reads only the folders whose files changed.
- **Measured.** The chat shows real token usage per turn. [`test_token_budget.py`](tests/integration/test_token_budget.py) fails if the fixed part goes over 4,000 tokens with 100 tasks or grows by more than 10 tokens per task.
Details: [docs/architecture.md → Scaling to 100+ tasks](docs/architecture.md#scaling-to-100-tasks).
---
## 6. Project tour
```
.env.example every setting, commented (copy to .env)
pyproject.toml / uv.lock dependencies and their exact locked versions; the 3 CLI commands
docs/
architecture.md flows, the approval gate, A2UI, scaling, errors
adding-a-task.md the partner guide (contract fields, views, size limits, checklist)
tasks/ ── DATA: partner task contracts (no code) ──
categories/ task.yaml, functions.yaml (8), ui.yaml (2 views), prompts/ (3), skills/ (1)
_template/ the contract: *.schema.json (generated) + a commented brands example
src/mcp_server/ ── THE MCP SERVER (uv run salla-mcp serve) ──
server.py entry point: wires engine + api; LLM instructions; CLI serve | validate | schema
config.py settings from .env
engine/ core logic: no MCP or HTTP code
contracts.py the partner contract (Pydantic): validation, request building, size limits, views
catalog.py live catalog: load, validate, BM25 search, install/uninstall, CatalogScope
runtime.py the Engine: call / execute / render / apply
plans.py change plans: staged writes + "$wN.id" placeholders
sandbox.py Monty sandbox + the typed stubs scripts are checked against
store_client.py HTTP to Salla: auth, retries, readable errors
a2ui.py A2UI: load-time checks of views + the messages that show them
a2ui_spec/ the official A2UI v0.9.1 JSON Schemas (copied, Apache-2.0)
api/ how clients reach the engine (the only MCP / HTTP code)
tools.py the 6 MCP tools
approval.py the approval question and its answer (MCP elicitation)
protocol.py resources, prompts, pagination, profile scoping (replaces SDK handlers)
admin.py admin HTTP API: validate / install / uninstall tasks live
profiles.py one endpoint per kind of agent (/mcp?profile=<name>)
src/chat_agent/ ── THE CHAT AGENT (uv run salla-agent) ──
main.py FastAPI + WebSocket; one MCP session per browser tab
mcp_connection.py connect by URL + paginated discovery
agent.py agent loop (OpenAI Responses API), MCP ⇄ OpenAI, A2UI hand-off, usage
hitl.py the HITL broker: pauses on any MCP elicitation, asks the browser, resumes with the answer
static/ index.html, app.js, style.css
src/task_studio/ ── THE TASK CONTRACT STUDIO (uv run salla-studio) ──
app.py the page + /api/* in front of the MCP admin API (token stays server-side)
files.py draft ⇄ contract files (YAML + Markdown)
static/ index.html, studio.js, studio.css
src/a2ui_renderer/ ── A2UI RENDERER, shared by the chat and the Studio ──
static/ a2ui.js (the renderer), a2ui.css
tests/ mcp_server/, agent/, studio/ (each app alone) + integration/ (apps together)
```
Rules the layout follows:
- **Isolated apps.** The three apps never import each other; they talk over HTTP only. Both web apps serve the shared A2UI renderer the way they'd use a CDN script.
- **Protocol-free engine.** `engine/` has no MCP or HTTP imports; `api/` is the only MCP code.
- **One place for SDK internals.** `api/protocol.py` is the only module that touches them.
| Concept | Where to look |
|---|---|
| Partner contract (Pydantic → JSON Schema) | [`engine/contracts.py`](src/mcp_server/engine/contracts.py), [`tasks/_template/`](tasks/_template/) |
| Live catalog, search, scopes | [`engine/catalog.py`](src/mcp_server/engine/catalog.py) |
| call / execute / render / apply | [`engine/runtime.py`](src/mcp_server/engine/runtime.py) |
| Change plans and placeholders | [`engine/plans.py`](src/mcp_server/engine/plans.py) |
| Code Mode sandbox + typed stubs | [`engine/sandbox.py`](src/mcp_server/engine/sandbox.py) |
| A2UI views | [`engine/a2ui.py`](src/mcp_server/engine/a2ui.py), [`tasks/categories/ui.yaml`](tasks/categories/ui.yaml), [`a2ui.js`](src/a2ui_renderer/static/a2ui.js) |
| The 6 tools | [`api/tools.py`](src/mcp_server/api/tools.py) |
| Human in the loop (elicitation) | [`api/approval.py`](src/mcp_server/api/approval.py), [`chat_agent/hitl.py`](src/chat_agent/hitl.py) |
| Resources, prompts, pagination, profiles | [`api/protocol.py`](src/mcp_server/api/protocol.py), [`api/profiles.py`](src/mcp_server/api/profiles.py) |
| Admin API | [`api/admin.py`](src/mcp_server/api/admin.py) |
| Agent loop, token usage | [`chat_agent/agent.py`](src/chat_agent/agent.py) |
---
## 7. Build a similar project from scratch
Use this as a recipe for your own API (Shopify, Zid, an internal API…). Each step names the file here to copy ideas from.
**Step 0. Set up the project.**
```bash
uv init my-mcp-demo && cd my-mcp-demo
uv add "mcp==2.2.*" httpx pydantic-settings "pydantic-monty==0.0.23" pyyaml jsonschema openai fastapi "uvicorn[standard]"
uv add --dev pytest pytest-asyncio respx
```
Use a `src/` layout with one package per app, and declare them in `pyproject.toml` under `[tool.hatch.build.targets.wheel] packages`. Define the CLI commands under `[project.scripts]`, and set `[tool.pytest.ini_options] asyncio_mode = "auto"`.
**Steps, in the order that works best:**
| # | Build | Key decisions | Look at |
|---|---|---|---|
| 1 | **The contract** | Pydantic models with `extra="forbid"`, so a typo is an error. Each function has `effect` (read/write/destructive), `method`, `path`, JSON-Schema `params` and an optional `request` mapping. Export JSON Schemas for partners' editors, and add size limits. | `engine/contracts.py`, `tasks/_template/` |
| 2 | **The catalog** | Strict loading: an invalid task stops the server and names the file and field. Live `install` / `uninstall`, re-reading only changed folders. Search over everything. | `engine/catalog.py` |
| 3 | **The HTTP executor** | Contract + arguments → request: path placeholders, query or body, dotted keys for nesting. Retries and readable errors. | `FunctionContract.build_request`, `engine/store_client.py` |
| 4 | **Change plans** | Writes are staged, never run. `$wN.id` placeholders connect writes within one plan. Apply in order, stop at the first failure, report every change. | `engine/plans.py`, `Engine.apply` |
| 5 | **The sandbox** | Monty: no filesystem, network or environment. Type-check scripts against generated stubs before running them, and hand the script only the functions it names. | `engine/sandbox.py`, `functions_used` |
| 6 | **The MCP tools** | A fixed set (search, describe, call, execute, render_ui, apply_changes), with read-only annotations. Engine errors become `ToolError`s the LLM can read. | `api/tools.py` |
| 7 | **The approval gate** | A resolver (`Resolve(confirm_plan)`) fills `apply_changes`' approval parameter through elicitation, so the LLM can't approve itself. It is deterministic, so retries ask the same question. | `api/tools.py`, `api/approval.py` |
| 8 | **Protocol extras** | Paginated `resources/list` and `prompts/list`, prompts served from Markdown, `resources/read`, and profile scoping, via `_lowlevel_server.add_request_handler`. | `api/protocol.py`, `api/profiles.py` |
| 9 | **Admin API** | `custom_route`s next to `/mcp`, protected by a bearer token (compared with `hmac`). They validate with the start-up rules, then install. | `api/admin.py` |
| 10 | **The agent** | `mcp.Client(url, elicitation_callback=…)`, paginated discovery, and MCP tool → OpenAI function conversion. The Responses API with `previous_response_id` and `parallel_tool_calls=False`, so one approval happens at a time. | `chat_agent/agent.py`, `mcp_connection.py` |
| 11 | **Human in the loop in the browser** | One WebSocket per tab. An `ApprovalBroker` turns elicitation into a card and waits for the answer. The agent loop runs as a background task, so answers can arrive. | `chat_agent/main.py`, `hitl.py`, `static/app.js` |
| 12 | **Skills on demand** | Skills are `skill://` resources, found with search and read by the agent's `load_skill` tool, never listed in the prompt. | `api/protocol.py`, `agent.py` |
| 13 | **A2UI views** | Copy the official v0.9.1 schemas, check views at load time with a `referencing` registry, and build the three messages. A tiny DOM-only renderer draws them, and clicks come back as `[UI action]`. | `engine/a2ui.py`, `src/a2ui_renderer/` |
| 14 | **The Studio** | A form to a draft to files, then the admin API. The token stays server-side. It uses the same renderer for previews. | `src/task_studio/` |
| 15 | **Scaling guards** | Size limits, BM25 search, profiles, clipped results, token usage, and a token-budget test with 100 generated tasks. | [Built for 100+ tasks](#5-built-for-100-tasks) |
| 16 | **Tests** | Run the MCP server in memory (`Client(server)`), mock the API with `respx`, script the LLM with a fake. They need no keys and run in about 20 s. | `tests/` |
**Design rules worth keeping:**
- Tasks are data. Adding one never changes code.
- The tool count is fixed, and the prompt must not grow with the tasks.
- Every change passes one approval gate, which fails closed.
- The LLM can't approve, invent screens or reach undeclared endpoints.
- The engine has no protocol code, and apps talk over HTTP only.
- Check everything at load time, with messages that name the file and field.
---
## 8. Tests
```bash
uv run pytest # everything (90 tests, about 20 s)
uv run pytest tests/mcp_server # the server alone
uv run pytest tests/integration # apps together (agent ⇄ server, Studio ⇄ server, token budget)
```
The server runs in memory, Salla's API is mocked with `respx`, the sandbox runs for real, and the LLM is scripted. Among other things, the tests check that:
- a **rejected plan sends no write request**, and a client that can't show approval forms can't change the store;
- placeholders become real IDs, and apply stops at the first failure;
- type errors reach the LLM before any code runs, and the sandbox has no filesystem and a time limit;
- contracts are strict, the template schemas match the engine, and texts respect the size limits;
- views are valid A2UI, and `render_ui` only accepts data that matches the view, including from inside scripts;
- profiles hide everything outside their tasks, and search ranks the best match first;
- with 100 tasks, the fixed part of the prompt stays within its token budget.
---
## 9. Troubleshooting
| Symptom | Fix |
|---|---|
| The agent says "Could not talk to the MCP server" | Start it with `uv run salla-mcp serve`, and check `MCP_SERVER_URL`. |
| The server exits with "SALLA_ACCESS_TOKEN is not set" | Create `.env` from `.env.example`, and start the app from the repository root. |
| The server exits with "Invalid task contract(s)" | Run `uv run salla-mcp validate`; the message names the file and the field. |
| `401` / `403` from Salla | The token has expired, or the app is missing the `categories.read_write` scope. |
| The LLM's script fails with "type errors" | Expected: it gets the message back and fixes the script. |
| The Studio says the admin API is disabled | Put the same `MCP_ADMIN_TOKEN` in `.env` for both processes and restart them. |
| `Unknown profile 'x'` | The `?profile=` in `MCP_SERVER_URL` must be a name from the server's `MCP_PROFILES`. |
| Windows: `uv sync` fails with "failed to remove file … (os error 32)" | A running app holds the `.exe` launcher. Stop the servers, then run `uv sync` again. |
| Windows: `UnicodeEncodeError` when printing Arabic | The console isn't UTF-8. Run `$env:PYTHONIOENCODING="utf-8"` (PowerShell) or `set PYTHONIOENCODING=utf-8` (cmd) first. |
| The chat shows raw Markdown | The browser couldn't load marked / DOMPurify from the CDN; check internet access. |
---
## 10. Versions, limits and history
**History** (git):
- tag `v1-per-tool`: the first design, with one Python-coded MCP tool per operation;
- the `v2` commit: contract-driven Code Mode, the chat agent and the Studio;
- the current version adds A2UI views and the scaling work in section 5.
**Known limits:**
- **`PlanStore` is in memory.** With several server replicas, swap it for a Redis-backed class with the same methods.
- **Monty is pre-1.0 (0.0.23, pinned).** All sandbox code lives in `engine/sandbox.py`.
- **Only part of A2UI is drawn.** The renderer supports Column, Row, List, Card, Text, Image, Divider, Button, TextField and `formatString`. The server refuses views that use anything else.
- **A2UI checks are slow at start-up.** With 100 tasks, start-up takes about 7 s, mostly checking views against the A2UI schemas.
- **Token counts in the budget test are estimates** (4 characters per token). The chat's usage line shows real numbers from the API.
- **Reads can't see staged writes.** Stage related writes together, and read after applying.
**A note on the word "task":** here, a *task* is a business domain folder (categories, brands, …). The MCP specification also has an experimental **Tasks** feature for long-running requests. That is a different thing, and the MCP Python SDK 2.2 used here doesn't implement it.
**License and credits:**
- The repository doesn't have a license file yet.
- The A2UI JSON Schemas in `src/mcp_server/engine/a2ui_spec/` are © the A2UI authors, Apache-2.0 (see [NOTICE.md](src/mcp_server/engine/a2ui_spec/NOTICE.md)).
- The sandbox is [Monty](https://github.com/pydantic/monty) by Pydantic.
- The MCP SDK is the official [modelcontextprotocol/python-sdk](https://github.com/modelcontextprotocol/python-sdk).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues