xgen-taskmaster
by jinsoo96
README.md
# xgen-taskmaster
Show a task once in the browser. Get one tool an agent can call.
[한국어](README.ko.md)
xgen-taskmaster records one human demonstration of a web task, finds the HTTP calls that produce the result,
explains every value those calls send (typed by the person, returned by an earlier call, or constant),
and saves the whole procedure as a declarative recipe. A single generic interpreter replays the recipe,
so the result is one function-calling tool or MCP tool per intent, with no site-specific code.
The core is standard library only. Recording uses Playwright (optional extra). A small LLM is used only to
name the tool and its parameters; the procedure itself is built and checked without any model.
## Pipeline
```
person shows the task once
│
▼
┌────────────────┐ HAR + typed values + clicks
│ 1. record │ ─────────────────────────────┐
└────────────────┘ ▼
┌──────────────────┐
│ 2. build │ exact string matching, no LLM
│ keep the calls │ (params, bindings, constants,
│ the result │ pruning, extraction rules)
│ depends on │
└────────┬─────────┘
▼
┌──────────────────┐
│ 3. label (opt.) │ small LLM names the tool and params
└────────┬─────────┘
▼
┌──────────────────┐
│ 4. verify │ replay: status class + response shape
└────────┬─────────┘ per step, new values pass, nonsense fails
▼
┌───────────────────────────┼─────────────────────────────┐
▼ ▼ ▼
function-calling tool MCP tool (stdio) stdin JSON -> stdout JSON script
```
## Value explanation rules
Every value a kept request sends (path segment, query value, header, form field, JSON leaf) is classified in this order.
| Order | Evidence | Becomes | Example |
|---|---|---|---|
| 1 | equals or contains a value the person typed | parameter (`params.x`), or secret (`secrets.x`) when typed into a login form | `query=김영희` -> `{{params.query}}` |
| 2 | equals an id or token in an earlier JSON response, header or Set-Cookie | binding with a JSON path | `Bearer {{steps.post_login.json.access_token}}` |
| 2a | the earlier value sits inside a list item that also holds a typed value | binding with a filter instead of an index | `items[?name=params.query].customer_id` |
| 3 | found inside an earlier HTML or text response | extraction rule (shortest stable left context) | CSRF token from `<meta name="csrf-token">` |
| 4 | one-off client value (UUID, current timestamp) seen nowhere else | regenerated per run | `{{gen.uuid}}` |
| 5 | none of the above | constant | `page=1` |
Only values that look generated (ids, tokens, hashes) take part in rules 2 and 3, so ordinary words and MIME types
are not mistaken for data flow. Unexplained values in authorization-like headers become secrets instead of being
stored. Calls that the result does not depend on (page loads, analytics beacons, profile polls) are dropped.
## Quick start
```bash
pip install "xgen-taskmaster[record]"
playwright install chromium
xgen-taskmaster record https://intranet.example.com/ -o loans.rec.json # do the task, close the window
xgen-taskmaster build loans.rec.json -o loans.recipe.json \
--llm-base-url http://vllm:8000/v1 --llm-model qwen3-32b \
--llm-extra '{"chat_template_kwargs": {"enable_thinking": false}}'
xgen-taskmaster verify loans.recipe.json --secrets secrets.json
echo '{"customer_name": "이철수"}' | xgen-taskmaster run loans.recipe.json --secrets secrets.json
xgen-taskmaster serve ./recipes --secrets secrets.json # MCP over stdio
```
`secrets.json` holds login values and is never written into a recipe:
`{"get_customer_loans": {"username": "...", "password": "..."}}`.
Python:
```python
from xgen_taskmaster import Recording, build, label, verify, run, function_tool, OpenAICompatible
recipe = build(Recording.load("loans.rec.json"))
recipe = label(recipe, OpenAICompatible(base_url, "qwen3-32b", extra_body={"enable_thinking": False}))
assert verify(recipe, {"customer_name": "이철수"}, secrets).ok
tool = function_tool(recipe) # pass to any OpenAI-compatible chat API
result = run(recipe, {"customer_name": "박민준"}, secrets)
```
## Results on the bundled demo sites
`examples/demo.py` reproduces everything below. Both sites issue new ids, tokens and ports on every start,
so replaying the recorded requests verbatim is rejected (401 on the loan desk) and only correct bindings succeed.
| | Loan desk | Order desk |
|---|---|---|
| Stack | REST JSON, bearer token | form login, session cookie, CSRF token in HTML, GraphQL |
| Task | look up a customer's loans | refund an order with a reason (write) |
| Requests recorded / steps kept | 6 / 3 | 6 / 4 |
| New values replayed correctly | 2 / 2 | 2 / 2 |
| Nonsense value rejected | yes | yes |
| Build time | 1.1 ms | 4.5 ms |
| Tool run (avg of 20) vs browser demo | 19 ms vs 1.5 s | 22 ms vs 1.4 s |
| Name from LLM (qwen3-32b) | `get_customer_loans(customer_name)` | `order_refund(email, reason)` |
Agent check with qwen3-32b and both tools available: 3 of 3 Korean requests ("박민준 고객 대출 잔액 좀 알려줘",
"이철수 고객님 대출 내역 보여줘", "junho@example.com 고객 주문 환불 처리해줘. 사유는 배송 지연이야")
were answered with exactly one correct tool call and a result equal to the server's own data.
The MCP server answered `tools/list` and `tools/call` correctly as a separate process.
## Recipe format
```json
{
"name": "get_customer_loans",
"origin": "http://127.0.0.1:8765",
"params": [{"name": "customer_name", "type": "string", "example": "김영희", "secret": false},
{"name": "password", "secret": true}],
"steps": [
{"id": "post_login", "method": "POST", "path": "/api/auth/login", "body_kind": "json",
"body": {"username": "{{secrets.username}}", "password": "{{secrets.password}}"}},
{"id": "get_customers", "method": "GET", "path": "/api/customers",
"query": [["query", "{{params.customer_name}}"], ["page", "1"]],
"headers": {"authorization": "Bearer {{steps.post_login.json.access_token}}", "x-trace-id": "{{gen.uuid}}"}},
{"id": "get_loans", "method": "GET",
"path": "/api/customers/{{steps.get_customers.json.items[?name=params.customer_name].customer_id}}/loans",
"headers": {"authorization": "Bearer {{steps.post_login.json.access_token}}"}}
],
"returns": "steps.get_loans.json"
}
```
References: `params.<name>`, `secrets.<name>`, `gen.uuid|now_s|now_ms`,
`steps.<id>.json|header|cookie|extract.<path>`. Paths support `a.b`, `[0]`, `["odd key"]` and `[?field=ref]`.
Cookies travel in a per-run cookie jar; redirects are not followed, so each recorded exchange stays one step.
## Limits
- One origin per recipe. Third-party hosts are ignored.
- Requests signed by page JavaScript, CAPTCHAs, bot checks and TLS fingerprinting are out of scope.
- Filtering done only in the browser cannot be replayed; typed values that never reach the network are listed in `meta.unused_inputs`.
- A clicked list item is matched by a typed value when one exists in the item, otherwise by its position.
- Writes are not idempotent. Verify write recipes against a test system.
## Related work
Traffic-level: Integuru (an LLM marks dynamic parts, then substring search; AGPL), Unbrowse (compares two recordings;
learning service not open), mitmproxy2swagger and har-to-openapi (OpenAPI specs, no data flow).
UI-level: browser-use workflow-use, Skyvern, Stagehand caches, WALT, SkillWeaver, Agent Workflow Memory.
xgen-taskmaster differs by deciding every value without a model and replaying HTTP calls instead of UI actions.
No code from AGPL projects is used.
## License
Source-available, not open source. See [LICENSE](LICENSE). PlateerLab organization members may use it within
Plateer products and services (Section 4).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues