Skip to main content
Glama

xgen-taskmaster

Show a task once in the browser. Get one tool an agent can call.

한국어

xgen-taskmaster records one human demonstration of a web task, finds the HTTP calls that produce the result, explains every value those calls send (typed by the person, returned by an earlier call, or constant), and saves the whole procedure as a declarative recipe. A single generic interpreter replays the recipe, so the result is one function-calling tool or MCP tool per intent, with no site-specific code.

The core is standard library only. Recording uses Playwright (optional extra). A small LLM is used only to name the tool and its parameters; the procedure itself is built and checked without any model.

Pipeline

 person shows the task once
          │
          ▼
 ┌────────────────┐   HAR + typed values + clicks
 │ 1. record      │ ─────────────────────────────┐
 └────────────────┘                              ▼
                                       ┌──────────────────┐
                                       │ 2. build         │  exact string matching, no LLM
                                       │  keep the calls  │  (params, bindings, constants,
                                       │  the result      │   pruning, extraction rules)
                                       │  depends on      │
                                       └────────┬─────────┘
                                                ▼
                                       ┌──────────────────┐
                                       │ 3. label (opt.)  │  small LLM names the tool and params
                                       └────────┬─────────┘
                                                ▼
                                       ┌──────────────────┐
                                       │ 4. verify        │  replay: status class + response shape
                                       └────────┬─────────┘  per step, new values pass, nonsense fails
                                                ▼
                    ┌───────────────────────────┼─────────────────────────────┐
                    ▼                           ▼                             ▼
          function-calling tool          MCP tool (stdio)          stdin JSON -> stdout JSON script

Related MCP server: Hydra ACI

Value explanation rules

Every value a kept request sends (path segment, query value, header, form field, JSON leaf) is classified in this order.

Order

Evidence

Becomes

Example

1

equals or contains a value the person typed

parameter (params.x), or secret (secrets.x) when typed into a login form

query=김영희 -> {{params.query}}

2

equals an id or token in an earlier JSON response, header or Set-Cookie

binding with a JSON path

Bearer {{steps.post_login.json.access_token}}

2a

the earlier value sits inside a list item that also holds a typed value

binding with a filter instead of an index

items[?name=params.query].customer_id

3

found inside an earlier HTML or text response

extraction rule (shortest stable left context)

CSRF token from <meta name="csrf-token">

4

one-off client value (UUID, current timestamp) seen nowhere else

regenerated per run

{{gen.uuid}}

5

none of the above

constant

page=1

Only values that look generated (ids, tokens, hashes) take part in rules 2 and 3, so ordinary words and MIME types are not mistaken for data flow. Unexplained values in authorization-like headers become secrets instead of being stored. Calls that the result does not depend on (page loads, analytics beacons, profile polls) are dropped.

Quick start

pip install "xgen-taskmaster[record]"
playwright install chromium

xgen-taskmaster record https://intranet.example.com/ -o loans.rec.json     # do the task, close the window
xgen-taskmaster build loans.rec.json -o loans.recipe.json \
    --llm-base-url http://vllm:8000/v1 --llm-model qwen3-32b \
    --llm-extra '{"chat_template_kwargs": {"enable_thinking": false}}'
xgen-taskmaster verify loans.recipe.json --secrets secrets.json
echo '{"customer_name": "이철수"}' | xgen-taskmaster run loans.recipe.json --secrets secrets.json
xgen-taskmaster serve ./recipes --secrets secrets.json                     # MCP over stdio

secrets.json holds login values and is never written into a recipe: {"get_customer_loans": {"username": "...", "password": "..."}}.

Python:

from xgen_taskmaster import Recording, build, label, verify, run, function_tool, OpenAICompatible

recipe = build(Recording.load("loans.rec.json"))
recipe = label(recipe, OpenAICompatible(base_url, "qwen3-32b", extra_body={"enable_thinking": False}))
assert verify(recipe, {"customer_name": "이철수"}, secrets).ok
tool = function_tool(recipe)            # pass to any OpenAI-compatible chat API
result = run(recipe, {"customer_name": "박민준"}, secrets)

Results on the bundled demo sites

examples/demo.py reproduces everything below. Both sites issue new ids, tokens and ports on every start, so replaying the recorded requests verbatim is rejected (401 on the loan desk) and only correct bindings succeed.

Loan desk

Order desk

Stack

REST JSON, bearer token

form login, session cookie, CSRF token in HTML, GraphQL

Task

look up a customer's loans

refund an order with a reason (write)

Requests recorded / steps kept

6 / 3

6 / 4

New values replayed correctly

2 / 2

2 / 2

Nonsense value rejected

yes

yes

Build time

1.1 ms

4.5 ms

Tool run (avg of 20) vs browser demo

19 ms vs 1.5 s

22 ms vs 1.4 s

Name from LLM (qwen3-32b)

get_customer_loans(customer_name)

order_refund(email, reason)

Agent check with qwen3-32b and both tools available: 3 of 3 Korean requests ("박민준 고객 대출 잔액 좀 알려줘", "이철수 고객님 대출 내역 보여줘", "junho@example.com 고객 주문 환불 처리해줘. 사유는 배송 지연이야") were answered with exactly one correct tool call and a result equal to the server's own data. The MCP server answered tools/list and tools/call correctly as a separate process.

Recipe format

{
 "name": "get_customer_loans",
 "origin": "http://127.0.0.1:8765",
 "params": [{"name": "customer_name", "type": "string", "example": "김영희", "secret": false},
            {"name": "password", "secret": true}],
 "steps": [
  {"id": "post_login", "method": "POST", "path": "/api/auth/login", "body_kind": "json",
   "body": {"username": "{{secrets.username}}", "password": "{{secrets.password}}"}},
  {"id": "get_customers", "method": "GET", "path": "/api/customers",
   "query": [["query", "{{params.customer_name}}"], ["page", "1"]],
   "headers": {"authorization": "Bearer {{steps.post_login.json.access_token}}", "x-trace-id": "{{gen.uuid}}"}},
  {"id": "get_loans", "method": "GET",
   "path": "/api/customers/{{steps.get_customers.json.items[?name=params.customer_name].customer_id}}/loans",
   "headers": {"authorization": "Bearer {{steps.post_login.json.access_token}}"}}
 ],
 "returns": "steps.get_loans.json"
}

References: params.<name>, secrets.<name>, gen.uuid|now_s|now_ms, steps.<id>.json|header|cookie|extract.<path>. Paths support a.b, [0], ["odd key"] and [?field=ref]. Cookies travel in a per-run cookie jar; redirects are not followed, so each recorded exchange stays one step.

Limits

  • One origin per recipe. Third-party hosts are ignored.

  • Requests signed by page JavaScript, CAPTCHAs, bot checks and TLS fingerprinting are out of scope.

  • Filtering done only in the browser cannot be replayed; typed values that never reach the network are listed in meta.unused_inputs.

  • A clicked list item is matched by a typed value when one exists in the item, otherwise by its position.

  • Writes are not idempotent. Verify write recipes against a test system.

Traffic-level: Integuru (an LLM marks dynamic parts, then substring search; AGPL), Unbrowse (compares two recordings; learning service not open), mitmproxy2swagger and har-to-openapi (OpenAPI specs, no data flow). UI-level: browser-use workflow-use, Skyvern, Stagehand caches, WALT, SkillWeaver, Agent Workflow Memory. xgen-taskmaster differs by deciding every value without a model and replaying HTTP calls instead of UI actions. No code from AGPL projects is used.

License

Source-available, not open source. See LICENSE. PlateerLab organization members may use it within Plateer products and services (Section 4).

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    F
    maintenance
    Enables creation of reusable browser automation skills through demonstration by recording user actions in a browser while narrating, then converting those workflows into executable skills that can be invoked through natural language.
    1
    -
  • -
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to discover and call business actions from browser-only legacy web applications as typed MCP tools, executing them through the original GUI via Playwright.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables agents to securely discover and invoke a centrally governed catalog of tools from distributed internal and external providers, with policy enforcement, quotas, inspection, and audit controls.
    -