AginxBrowser
<p align="center"><img src="web/brand/icon.svg" width="160" alt="AginxBrowser"></p>
# AginxBrowser
English | [中文](README.zh-CN.md)
**The Browser for AI Agents. See the live web. Read it. Act on it. Remember it.**
[](LICENSE)
[](https://skills.sh/yinnho/aginxbrowser)
[](https://x.com/aginxbrowser)
A browser built for agents from the first line of code — not a human browser bolted onto automation. See the world, read it, search it, act on it, and keep what you read: one Rust binary with built-in V8, **no Chromium required**.
> Humans have Chrome. Agents have AginxBrowser.
One binary, zero dependencies, instant service. The HTTP API is the whole interface — agents plug in and go.
<video src="https://github.com/yinnho/aginxbrowser/releases/download/v0.5.3/lightpanda-star-story.mp4" controls muted width="720"></video>
*The star that got our attention: Pierre Tachoire, co-founder of [Lightpanda](https://lightpanda.com) — the headless browser our [bench](bench/README.md) measures against — starred the repo. 90 seconds on why that mattered to us.*
*Real pages rendered by AginxBrowser's diting engine (no Chromium) — Wikipedia, this repo, Rust. [Screenshot it yourself →](docs/API.md#screenshot)*

## Why Agents Need Their Own Browser
Measured against headless Chrome on the same 20 pages, same network ([bench](bench/README.md), 2026-08-28): **7.6× faster** to agent-usable text (p50 532 ms vs 4 053 ms), **~10× less memory** (227 MB for the whole process vs ~2.1 GB per Chrome page), and 0 hard failures where Chrome's `--dump-dom` produced no DOM on 5 of 40 loads. Re-run on v0.5.21 (2026-09-29) on a degraded-network day held the ratio — 5.4× p50, 468 MB whole-run vs 1.75 GB per page — both raw TSVs are committed. An agent's total cost is browser efficiency × model efficiency — this is the browser half.
Existing "browser automation" was built for humans or for one-shot scraping — not for agents:
| | AginxBrowser | Puppeteer/Playwright | Firecrawl | Browser-use |
|---|---|---|---|---|
| Designed for | **Agents first** | Human debugging | Scraping service | LLM wrapper |
| Dependencies | Single binary, no Chromium | Chromium ~500MB | Docker ~1GB | Chromium |
| Sees (screenshots) | ✅ built-in diting rendering engine | Needs Chromium | ❌ | Needs Chromium |
| Reads | markdown + js_extract + fetch receipts | DIY | markdown | DIY |
| Writes documents | ✅ `render_markdown`: deterministic HTML + inline-SVG diagrams | ❌ | ❌ | ❌ |
| Finds (search) | ✅ 20 engines, 7 categories, merged | ❌ | ❌ | ❌ |
| Acts | indexed session interaction | DevTools API | ❌ | LLM-driven |
| Remembers | ✅ local fetch/search cache (SQLite FTS5) | ❌ | crawl cache | ❌ |
| Protocol | HTTP | Node API | HTTP | Python |
| TLS fingerprints | ✅ Chrome/Firefox/Safari/Edge | Plugin required | ❌ | ❌ |
| CAPTCHA | ✅ detect + auto-wait + optional 2captcha | DIY | ❌ | ❌ |
| Interactive sessions | ✅ persistent | ✅ | ❌ | ✅ |
Same-tier engines, not the tools in the table above. Cells are capabilities, not speed.
| | AginxBrowser | Obscura | Blitz | Lightpanda |
|---|---|---|---|---|
| What it is | Rust browser, V8, diting CSS paint | Rust headless browser, V8 | HTML/CSS engine (Stylo). Not an agent browser | Zig headless browser, V8 |
| Screenshot | built-in paint, opt-in build | screenshots, screencast, PDF | paints a window | Hermes integration falls back to Chrome for screenshots |
| Public CSS suite | none in CI | not published as WPT | WPT in CI, including SVG | not published as WPT |
| License | Apache-2.0 | Apache-2.0 | Apache-2.0 and MIT | AGPL-3.0 |
An agent needs five things from a browser: **see, read, find, act, remember.** One binary covers them all — systemd-friendly, zero dependencies.
**Core advantage: no Chromium.** AginxBrowser inlines a full browser engine (V8 + Rust HTTP stack + the diting CSS/layout/paint rendering engine, with the Blitz/Stylo/Taffy lineage as its reference implementation). No Puppeteer, no Chrome, no Docker. One Rust binary under systemd is your agent browsing infrastructure.
## Two Things Stateless Renderers Can't Do
Most new "agent browsers" are stateless, fingerprint-less one-shot renderers — fine for public pages, dead on arrival against Cloudflare or login flows. AginxBrowser goes the opposite way:
- **🔐 Real TLS fingerprints** — stealth mode replicates the complete Chrome145 / Firefox133 / Safari / Edge TLS handshakes via BoringSSL (not just a UA string), switchable per request; Cloudflare Turnstile challenges wait automatically for `cf_clearance`. Fingerprint-less engines eat 403s — we get through.
- **🤝 Stateful interactive sessions** — login state injectable and exportable (`session_create(cookies=...)` ↔ `session_cookies`), surviving pagination and multi-step flows; `persistent: true` even survives idle eviction and server restarts — the same session id comes back logged in. One-shot engines throw state away.
> Reference point: Cloudflare's Kitesurf explicitly ships neither real TLS-fingerprint negotiation nor persistent auth sessions — anti-bot and login territory is exactly where AginxBrowser plays.
Apache-2.0 open source, single binary — self-host today, no cloud lock-in.
## Every Fetch Is a Receipt
Agents act on what a browser tells them, so the response reports what actually happened — not just "got a 200":
- **`tier`** — which path served the page: plain HTTP (~100 ms) or the V8-rendered browser tier. An agent can see *why* a fetch was fast or slow.
- **`redirected_from`** — the full redirect trail. `redirected_from[0]` is the URL you asked for, `url` is where the content actually came from — requested paired with effective, every hop visible.
- **`content_hash` + `changed_since_prev`** — every fetch is hashed; consecutive samples of the same URL can be diffed. A rate-limited origin serving the same frozen 200 body for days reads as `changed_since_prev: false` — the cheapest drift detector there is.
- **`captcha_event`** — when a challenge page was detected (and solved, if a solver is configured), the response says so instead of handing over a challenge page as if it were content.
The [local cache](#capabilities) builds on the same idea: search hits come back with `[§ heading]` section prefixes so an agent knows *where on the page* a hit landed, and ranking fuses keyword relevance with freshness.
## Capabilities
- **Tiered rendering**: static pages over plain HTTP (~100ms); V8 spins up only when JS rendering is needed (~1-2s) — 90% of the [bench](bench/README.md) page set served without spinning up V8 at all; every response reports which tier served it (`tier` field)
- **Multi-engine meta-search**: general web (Baidu / Bing / Sogou / WeChat / DuckDuckGo / Wikipedia / Hacker News), news (Bing News), code (Stack Overflow, GitHub, MDN), packages (npm, PyPI, RubyGems), academic (arXiv, OpenAlex), AI models (Hugging Face) — 20 engines across 7 categories, queried concurrently, merged and deduplicated. Operators can plug a private Meilisearch index into the same `/search`. Search → read in one step
- **Image search**: `categories=images` hits Baidu/Bing image indexes and returns direct binary `image_url` links (downloadable straight to jpg/png) plus `source_url` provenance
- **Interactive sessions**: persistent browser sessions with indexed interaction (`state/click/input/scroll/eval`) — agents browse like humans do, and `session_export` turns what an agent figured out into a runnable curl replay script (zero model tokens on re-run) — or, with `format=json`, into a flow document (`flow_run` replays it server-side with `{{var}}` substitution, `wait`/`expect` gates and saved outputs; installed flows live in `workflow/<name>/flow.json`, dropped in without a rebuild). Session tools also cover the acting part: `session_viewport` simulates device viewports (media queries respond), `session_wait` blocks on a selector or predicate with a timeout, `session_screenshot` renders the live state, `session_console` replays the page's console ring, and `session_storage` exports/restores cookies plus localStorage for login hand-off
- **Playback-link sniffer**: `session_network(filter=media)` extracts the m3u8/mp4/dash URLs a page's player *actually requested* at runtime — links found only in page HTML are often decoys, so the request log is the source of truth. `GET /session/{id}/har` exports the same traffic as HAR 1.2 (retained bodies included)
- **File download**: streaming to disk (no memory buffering), SHA-256 integrity, resume of interrupted transfers — for binaries, archives, datasets
- **Local cache that remembers**: every fetch/search lands in SQLite (FTS5) at `~/.aginxbrowser/cache.db` — a re-fetch inside the TTL answers from what the agent already read instead of re-paying network time, with CJK substring matching, `[§ heading]` section-aware snippets, per-URL content hashes for drift detection, and per-session scoping for shared deployments
- **CAPTCHA handling**: type detection with automatic Cloudflare challenge wait and optional 2captcha integration — search never stalls on verification pages
- **JS data extraction**: `js_extract` pulls `window.__INITIAL_STATE__` and other structured data out of SPAs
- **Document generation**: `POST /render_markdown` turns markdown into a deterministic, self-contained HTML artifact — the document layer, so agents never write HTML by hand. Prose rides a plain offline shell (no fonts, no scripts); fenced `archify` blocks carry typed zero-coordinate diagram JSON (sequence / workflow / architecture / dataflow / lifecycle families) and render to inline SVG via the layout engine. Same input, same bytes — the receipt carries the sha256 so determinism is verifiable. `theme` (light/dark) and `preset` (classic / signal-flow / blueprint / editorial) bake colors at generation time; `quality: "showcase"` is the delivery gate, grading route crossings, label clearance and rhythm without touching the artifact bytes. Guided-view tabs plus `window.agxViewer` (`focus` / ego / `route` / `reach`) make the artifact interactive. Mermaid sources are the agent's job to translate into archify JSON, not the engine's. Diagram vocabulary adapted from archify (MIT)
- **Screenshot rendering**: `/screenshot` endpoint (opt-in `--features screenshot`) paints the JS-rendered DOM with the diting rendering engine — pure CPU, no Chromium — to PNG. Vision input for agents
- **Timeline video**: `/video` renders a page's animation timelines to MP4 — the page's scripts register GSAP-style timelines in `window.__timelines` (`duration()` + `pause(t)`), each frame seeks to `t=i/fps` and paints the viewport, and the frames pipe into ffmpeg (H.264, yuv420p). Deterministic by construction: no wall clock in the pixel values, same render twice = same MP4. Needs ffmpeg on PATH
- **Page set (PDF/PNG/PPTX/DOCX)**: `/pdf` cuts a rendered page into pages and packages them — print mode paginates at top-level block boundaries (default A4 @96dpi, no half-cut text where a break can land on a block edge), slides mode makes one page per CSS-selector match sized to the element (an HTML deck with one `.slide` per page exports as a real deck). Image-based PDF: per-page JPEG via DCTDecode, hand-rolled PDF 1.4 writer, zero new dependencies. PPTX packages the same pages as one slide per page; DOCX as one page-sized section per page — both hand-rolled OOXML (stored-ZIP writer, fixed timestamps), byte-deterministic, zero new dependencies
- **TLS fingerprint spoofing**: stealth mode impersonates Chrome145/Firefox133/Safari/Edge, switchable per request
- **Firecrawl compatible**: `/v1/scrape` endpoint — existing Firecrawl clients migrate by changing the base URL
- **DNS rebinding protection**: built-in SSRF guard + post-resolution IP validation
## A Browser, Not a Crawler
AginxBrowser exists for **real-time retrieval**: an agent arrives with a question, reads a handful of pages, leaves with the answer. It is not a crawling tool — and the product is shaped so it can't quietly become one:
- **robots.txt is not our gate.** The RFC 9309 checker ships built in, but a real-time lookup layer isn't a crawler and doesn't do crawler etiquette by default; operators who want it set `AGINXBROWSER_HONOR_ROBOTS=1`.
- **No site-walking API.** There is no crawl endpoint and no link-following recursion — every page load happens because an agent asked for that page.
- **Built-in budgets.** Per-domain: 20 pages/minute. Per interactive session: 200 pages. Toggled via `AGINXBROWSER_DOMAIN_RATE_PER_MIN` / `AGINXBROWSER_SESSION_PAGE_LIMIT` (`0` disables on your own instance). Generous for an agent grinding through docs or a console; fatal to the page-after-page crawl pattern, including subdomain rotation (one registrable domain, one budget).
- **The hosted instance (browser.aginx.net) runs tighter budgets.** Every user shares one egress IP, and keeping sites comfortable with that IP is part of the service. Self-host if you want different numbers.
- Need to bulk-crawl a site? Use a crawler. This isn't one, and it won't become one.
## What It's For
Not demos — real jobs agent browsers are doing today:
- **Grind through admin consoles** — AWS / App Store Connect / Google Play, dozens of menu layers per task. Let the agent click; it comes back only when authorization is needed.
- **Batch actions behind login** — fill carts, dig through order history, check pages that only render while logged in. Inject cookies, operate, export for reuse.
- **Past anti-bot walls** — Cloudflare protection, Turnstile challenges, TLS fingerprint checks. Stealth mode pushes through instead of retreating at 403.
- **The Chinese internet** — Baidu / Sogou / WeChat meta-search across 5 engines, correct Chinese page rendering. Not English-web-only.
- **On-the-spot scripting** — agent reads the page, writes JS, evals it: highlighted comparison tables, reflowed content, product filters on hidden parameters. GreaseMonkey-on-steroids.
- **Multimodal vision** — screenshots as visual input for look-and-judge flows: picking seats, recognizing layouts, verifying rendering.
## Where It Sits in the Computer-Use Stack
Computer-use agents come in two layers. **GUI-layer** stacks (Cua, desktop CUA agents) drive a whole machine: screenshots of a display in, X11 mouse/keyboard events out, a VM or container per session. **Engine-layer** browsers skip the desktop — the page itself is the machine. AginxBrowser is the engine layer:
| | GUI layer (desktop CUA) | AginxBrowser (engine layer) |
|---|---|---|
| Action space | screen pixels → OS input events | DOM: click by coordinates or selector, real event dispatch |
| State readout | screenshot of a display | structured DOM / accessibility tree, screenshot optional |
| Per-session cost | Xvfb + VM/container | one process, no display server |
| Frames | flattened into pixels | first-class: hit-testing descends into iframes, frame-scoped evaluate |
| Replay | best-effort (compositor timing) | action logs export as runnable curl / flow JSON |
## Quick Start
**One-command full install** (SKILL.md trigger surface + verification):
```bash
# Download -> inspect the contents -> run only after review (never blind-run network scripts)
curl -fsSL https://raw.githubusercontent.com/yinnho/aginxbrowser/main/skill.sh -o skill.sh
less skill.sh
bash skill.sh
```
**Install the skill trigger surface via [skills.sh](https://www.skills.sh)**:
```bash
npx skills add yinnho/aginxbrowser
```
**Or just hand this block to your AI agent** — it sets itself up:
```text
Set up aginxbrowser as your web access layer:
1. Install it locally and verify it is healthy:
brew install yinnho/aginxbrowser/aginxbrowser
aginxbrowser doctor # then start: aginxbrowser (listens on 0.0.0.0:8089)
curl http://127.0.0.1:8089/health
2. From now on, route web tasks through http://127.0.0.1:8089 instead of raw curl:
- read a page / scrape a URL -> POST /fetch {"url":...} (markdown output, Cloudflare bypassed by default)
- search the web -> POST /search {"q":...,"fetch_top":3}
- see a page -> POST /screenshot {"url":...}
- login / form / click-through-> POST /session/create -> /session/{id}/state -> /input /click -> /close
```
Self-hosting:
```bash
# macOS / Linux via Homebrew
brew install yinnho/aginxbrowser/aginxbrowser
aginxbrowser doctor # features + fonts + egress self-check
# Docker (Docker Hub, mirrored on GHCR)
docker run -p 8089:8089 yinnho/aginxbrowser:latest
# (or ghcr.io/yinnho/aginxbrowser:latest)
# Or the prebuilt binary (platform detect + sha256 + mirror fallback + doctor self-check)
# macOS / Linux / Windows (git-bash; prebuilt Windows ships from v0.3.1, full `stealth`+`screenshot` feature set from v0.4.0)
# Cautious: download -> inspect -> run (never blind-run network scripts)
curl -fsSL https://browser.aginx.net/install.sh -o install.sh
less install.sh && bash install.sh
# Or straight in, if you trust the repo:
# curl -fsSL https://browser.aginx.net/install.sh | sh
# GitHub slow/blocked? AGINXBROWSER_GH_PROXY=https://ghfast.top/ bash install.sh
aginxbrowser doctor # features + fonts + egress self-check
# Or build from source (--features stealth,screenshot or you lose both)
cargo build --release --features stealth,screenshot
# Start the service
./target/release/aginxbrowser
# → Listening on 0.0.0.0:8089
# Verify
curl http://127.0.0.1:8089/health
# → {"status":"ok","engine":"diting"}
# Fetch a page
curl -sS -X POST http://127.0.0.1:8089/fetch \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com"}'
# Search (snippets only by default; fetch_top grabs page bodies for the top N)
curl -sS -X POST http://127.0.0.1:8089/search \
-H "Content-Type: application/json" \
-d '{"q":"macbook price","max_results":5,"fetch_top":2,"max_chars_per":2000}'
# Create an interactive session
curl -sS -X POST http://127.0.0.1:8089/session/create \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com"}'
# → {"session_id":"s_1","url":"https://example.com/"}
```
## REST Routes
Every capability is plain HTTP — no SDK required. There is no `/openapi.json` (the routes are few enough to list here); full request/response fields for each route are in [docs/API.md](docs/API.md).
| Method | Path | What it does |
|---|---|---|
| GET | `/health` | Liveness + build commit, UA, TLS, capabilities |
| GET | `/doctor` | Deep self-check: engines live, fonts, egress |
| GET | `/engines` | Search engine catalog with live suspension state |
| POST | `/fetch` | Fetch a page → text/markdown/html. Params: `render_tier` (`auto`/`http`/`browser`), `max_chars`, `capture_xhr`, … |
| POST | `/search` | Multi-engine meta-search. Params: `engines`, `categories`, `max_results`, `time_range` (`day`/`week`/`month`/`year`), `fetch_top`, `max_chars_per`, `wait_secs` |
| POST | `/click` | Click a CSS selector on a page |
| POST | `/eval` | Evaluate JavaScript on a page |
| POST | `/download` | Stream a file to disk (sha256, resume) |
| POST | `/screenshot` | Render page → PNG (`screenshot` feature) |
| POST | `/video` | Render animation timelines → MP4 (`screenshot` feature) |
| POST | `/pdf` | Paginate page → PDF/PNG/PPTX/DOCX (`screenshot` feature) |
| POST | `/v1/scrape` | Firecrawl-compatible scrape (+ `actions`) |
| POST | `/flow/run` | Run a recorded/edited flow JSON to completion — zero model tokens (`name` runs an installed `workflow/<name>/flow.json`, `flow` is inline, `vars` substitute `{{placeholders}}`, `session_id` composes with imported login state) |
| POST | `/session/create` | Start an interactive session (cookies/UA carried across calls) |
| POST | `/import/curl` | Create a logged-in session from a DevTools "Copy as cURL" |
| GET | `/session/list` | Live sessions |
| POST | `/session/{id}/navigate` · `/state` · `/click` · `/click_xy` · `/drag` · `/input` · `/scroll` · `/eval` · `/wait` · `/dialog` · `/viewport` · `/screenshot` · `/clone` · `/close` | Session actions |
| GET | `/session/{id}/cookies` · `/storage` · `/console` · `/network` · `/har` · `/export` | Session inspection |
| POST | `/render_markdown` | Markdown → deterministic self-contained HTML artifact (+ inline-SVG diagrams) |
Every capability is one POST away — no SDK, no protocol adapter.
## Project Layout
```
aginxbrowser/
├── Cargo.toml
├── build.rs # V8 snapshot generation
├── js/
│ └── bootstrap.js # V8 bootstrap script
├── workflow/ # Flow assets: <name>/flow.json replayed by flow_run (drop-in, no rebuild)
├── README.md
├── docs/
│ └── API.md # Full API reference (HTTP)
├── bench/ # Benchmark harness + results (vs headless Chrome)
│ ├── README.md # methodology + numbers
│ ├── pages.txt # fixed 20-page set
│ ├── run.py # harness
│ ├── summarize.py # TSV → results table
│ └── results/ # raw run data
└── src/
├── main.rs # HTTP service entry & routing
├── server.rs # Business layer (fetch/click/eval/search)
├── session.rs # Interactive browser sessions
├── docgen/ # Document layer: markdown → deterministic HTML + inline-SVG diagrams
├── render.rs # Tiered rendering (HTTP direct → diting browser engine)
├── store.rs # Local fetch/search cache (SQLite FTS5, drift hashes)
├── download.rs # Streaming file download (sha256, resume)
├── robots.rs # RFC 9309 robots.txt checker (opt-in gate)
├── rate.rs # Per-domain + per-session budgets
├── captcha.rs # CAPTCHA detection & auto-solve
├── firecrawl_compat.rs # Firecrawl-compatible /v1/scrape endpoint
├── video.rs # Timeline video pump (__timelines seek → ffmpeg → MP4)
├── pages.rs # Page pump (print/slides pagination → PDF/PNG)
├── ooxml.rs # OOXML containers (image-based PPTX/DOCX, stored-ZIP writer)
├── doctor_cli.rs # `aginxbrowser doctor` self-check
├── browser.rs # Top-level API: Browser, BrowserBuilder
├── page.rs # Top-level API: Page, Element
├── config.rs # BrowserConfig
├── cookie.rs # CookieStore
├── error.rs # Error types
├── search/ # 20 native search engines, 7 categories
│ ├── mod.rs # SearchEngine trait, Registry, merge/dedupe, progressive backoff
│ ├── baidu.rs # Baidu (JSON API, wreq stealth)
│ ├── baidu_images.rs # Baidu Images (acjson API, images category)
│ ├── bing.rs # Bing (HTML parsing, plain reqwest)
│ ├── bing_images.rs # Bing Images (images/async endpoint, images category)
│ ├── bing_news.rs # Bing News infinite-scroll fragment (news category; direct-first/proxy-retry)
│ ├── sogou.rs # Sogou web (HTML parsing, plain reqwest)
│ ├── sogou_wechat.rs # Sogou WeChat (HTML parsing + /link resolution)
│ ├── duckduckgo.rs # DuckDuckGo (html.duckduckgo.com, general; direct-first)
│ ├── wikipedia.rs # Wikipedia (MediaWiki search API, general; direct-first/proxy-retry)
│ ├── hn.rs # Hacker News (Algolia API, general; time_range filters created_at)
│ ├── stackexchange.rs # Stack Overflow (SE API v2.3, code category)
│ ├── mdn.rs # MDN Web Docs (v1 search API, code category only)
│ ├── github_repos.rs # GitHub repos (api.github.com, code category)
│ ├── arxiv.rs # arXiv (Atom API, academic category)
│ ├── openalex.rs # OpenAlex works (academic; DOI links, inverted-index abstracts)
│ ├── huggingface.rs # HF Hub models/datasets/spaces (ai category)
│ ├── npm.rs # npm packages (npms.io API, packages category)
│ ├── pypi.rs # PyPI name resolution (JSON API, packages)
│ ├── rubygems.rs # RubyGems gems (packages; direct-first/proxy-retry)
│ └── meilisearch.rs # Private-index adapter (env-configured)
│
├── diting_dom/ # HTML parsing, DOM tree, CSS selectors
├── diting_css/ # CSS parsing + cascade
├── diting_net/ # HTTP client, cookies, encoding, proxies
├── diting_js/ # V8 runtime, JS ops, module loading
├── diting_layout/ # Taffy-based layout, floats, hit-testing
├── diting_fonts/ # Bundled CJK font subset, fallback
└── diting_browser/ # Page navigation, lifecycle, browser context
```
## Build
```bash
# Standard build (no stealth; TLS fingerprint features inactive)
cargo build --release
# With stealth (requires go + cmake + C++ toolchain; enables TLS fingerprint spoofing)
cargo build --release --features stealth
# With screenshot rendering (enables /screenshot; adds the rendering stack, +30-40MB)
cargo build --release --features screenshot
# Full featured (recommended for production)
cargo build --release --features stealth,screenshot
```
Requirements: Rust 1.78+; the V8 static library downloads automatically on first build. The stealth feature additionally needs `go`, `cmake`, and a C++ compiler. The screenshot feature ships with a bundled CJK font subset (GB2312 + common symbols) — no system fonts required for correct Chinese rendering.
If your network can't reach the rusty_v8 CDN (build hangs with zero progress after "downloading v8"), pre-fill `~/.cache/rusty_v8` with the `librusty_v8.a.gz` for your version (fetch it from any reachable mirror/host and gunzip into place) and the build script skips the download.
## Runtime Environment Variables
| Variable | Default | Description |
|------|------|------|
| `AGINXBROWSER_BIND` | `0.0.0.0:8089` | Listen address |
| `AGINXBROWSER_STEALTH` | enabled | `0` disables stealth (for diagnostics) |
| `AGINXBROWSER_UA` | Linux Chrome145 | Spoofed User-Agent |
| `AGINXBROWSER_ACCEPT_LANGUAGE` | `zh-CN,zh;q=0.9,en;q=0.8` | Accept-Language header |
| `AGINXBROWSER_PROXY` | none | Optional fallback proxy. Blocked-source engines (Wikipedia, Bing News, Hugging Face, RubyGems) connect directly first and fall through to this proxy only when the direct attempt fails — overseas deployments need no proxy at all; per-request `use_proxy:true` also routes fetch/search through it. Browser/session navigations to known-blocked domains (wikipedia.org, github.com, …) route through it automatically. Standard `HTTP_PROXY`/`HTTPS_PROXY`/`ALL_PROXY` are deliberately ignored by the engine (set them for other tools freely); startup logs a warning when it sees one |
| `AGINXBROWSER_NAV_CHAIN_LIMIT` | `10` | JS navigation-chain cap: documents a page may chain via `location`/form hops before navigation aborts. The count includes the requested document (10 = initial doc + 9 hops). Raise for legit long chains (SSO handover across providers); HTTP 3xx redirects are budgeted separately (20, per Fetch spec / browser parity) |
| `AGINXBROWSER_CACHE_TTL_SECS` | `600` | `/fetch` cache TTL, `0` disables |
| `AGINXBROWSER_HONOR_ROBOTS` | unset | robots.txt is not consulted by default on `/fetch`, `/screenshot`, `/download`; set `1` to opt in (operator choice) |
| `AGINXBROWSER_ALLOW_FILE_ACCESS` | unset | Opt in to `file://` reads — navigation, subresources, `/fetch`. Same as the `--allow-file-access` CLI flag. Off by default: the server binds 0.0.0.0, so an open gate hands local files to anyone who can reach the port. Set it on a local dev instance, not a hosted one |
| `AGINXBROWSER_ALLOW_PRIVATE_NETWORK` | unset | Opt in to loopback/RFC1918/link-local fetches (the SSRF gate). Same as the `--allow-private-network` CLI flag — dev machines only |
| `AGINXBROWSER_ALLOW_NETWORK` | unset | Scoped alternative: comma-separated CIDR allowlist (e.g. `10.20.0.0/16,192.168.1.0/24`) that opens just those ranges — cloud-metadata endpoints (169.254.169.254, 100.100.100.200) and everything else stay blocked. Same as `--allow-network <cidrs>` |
| `AGINXBROWSER_FONT_DIR` | unset | Directory of extra fonts (`.ttf`/`.otf`/`.ttc`) loaded as tail fallbacks for scripts the bundled CJK subset doesn't cover (Korean, Thai, Arabic, …). Same as the `--font-dir <path>` CLI flag. Faces the bundle already covers keep bundle rendering — dir fonts are coverage tails, not named-family overrides |
| `AGINXBROWSER_ROBOTS_TTL_SECS` | `3600` | Per-host robots.txt policy cache TTL |
| `AGINXBROWSER_DOMAIN_RATE_PER_MIN` | `20` | Per-registrable-domain page budget per minute (subdomains share one budget); over-budget requests get 429 with the stance message. `0` disables. See "A Browser, Not a Crawler" |
| `AGINXBROWSER_SESSION_PAGE_LIMIT` | `200` | Total pages one interactive session may walk (navigation-causing clicks count); over-budget navigations are refused, the current page stays interactive. `0` disables |
| `AGINXBROWSER_STORE` | on | Local fetch/search cache; `0`/`false`/`off` disables |
| `AGINXBROWSER_STORE_PATH` | `~/.aginxbrowser/cache.db` | SQLite database location (created 0600) |
| `AGINXBROWSER_STORE_TTL_HOURS` | `720` | Cached page TTL |
| `AGINXBROWSER_STORE_SEARCH_TTL_HOURS` | `168` | Cached search-result-set TTL |
| `AGINXBROWSER_STORE_SCOPE` | `global` | `session` scopes the cache per session instead of one shared pool |
| `CAPTCHA_SOLVER_API_KEY` | none | 2captcha API key; enables CAPTCHA auto-solving |
| `CAPTCHA_SOLVER_SERVICE` | `2captcha` | CAPTCHA solving provider |
| `AGINXBROWSER_MEILI_URL` | none | Meilisearch base URL; set to enable the private-index engine |
| `AGINXBROWSER_MEILI_INDEX` | none | Meilisearch index uid to query |
| `AGINXBROWSER_MEILI_KEY` | none | Optional Bearer key for the Meilisearch instance |
## API Documentation
**Full API reference** → [`docs/API.md`](docs/API.md)
**Security audit notes** → [`docs/skills-sh-audit.md`](docs/skills-sh-audit.md) — why skills.sh shows "Critical Risk", and which real product feature each warning corresponds to
Covers:
- All HTTP endpoints (`/fetch`, `/search`, `/screenshot`, `/video`, `/pdf`, `/download`, `/render_markdown`, `/v1/scrape`, `/flow/run`, `/doctor`, the session endpoints)
- Environment variables, error codes, per-site scraping examples
## Plugging Into Other Systems
AginxBrowser is **pure attach-alongside infrastructure** — like a real browser, it runs as an independent service that anything can call, without embedding host code or polluting host config. Deploy one instance per machine (under systemd) and every app needing "render + scrape" capability shares it.
The attach point is HTTP — `/fetch`, `/search`, `/screenshot`, `/download`, `/render_markdown` for any language with an HTTP client.
Integration: read the environment variable `AGINXBROWSER_URL=http://127.0.0.1:8089`. Unset → behavior unchanged; set → risk-controlled sites automatically route through AginxBrowser for rendering, falling back gracefully on failure.
## Known Limitations
1. **Screenshots are opt-in**: `/screenshot` requires `cargo build --release --features screenshot` (adds the diting rendering stack). The default (and only) render engine in that build is diting — our own CSS+layout+paint stack, zero Blitz/Stylo code. The pinned-rev Blitz reference pipeline is a separate opt-in, `--features blitz-reference`, for comparison renders and the dual-engine cross-check tests. Complex-site CSS is approximate on both (not pixel-perfect like Chromium)
2. **Element coordinates supported**: `/screenshot` with `selector` returns element page coordinates (`selector_rects`, CSS px); `selector` alone crops directly to that element. Inline elements (`<a>text</a>`) get a rect too on the default diting engine — a union of their flattened inline content, strut-expanded to the element's own `line-height` like Chrome reports for replaced-only inlines (`<a><img></a>` → line-box height, not the image height). Empty inlines still have no rect — pick a block ancestor there
3. **JS interaction broadly works; heavy-fingerprint pages may still fail**: React/Vue event delegation works normally (URL-reflection attributes like `src`/`href` resolve to absolute URLs so Next.js/webpack hydrate and clicks trigger handlers). Heavy-fingerprint auth pages (WorkOS/Cloudflare) probing `navigator.plugins`, WebGL canvas etc. may still break until stealth fingerprint coverage completes
4. **Proxy support**: HTTP/HTTPS/SOCKS5 via `AGINXBROWSER_PROXY`
5. **Hard risk-controlled sites**: Baidu Wenku unsupported; Zhihu articles need a valid `__zse_ck`
## Star History
If AginxBrowser saved you a headless-Chrome fleet or a scraping headache, a star is how other agents (and their humans) find the project.
<a href="https://star-history.com/#yinnho/aginxbrowser&Date">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/svg?repos=yinnho/aginxbrowser&type=Date&theme=dark" />
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/svg?repos=yinnho/aginxbrowser&type=Date" />
<img alt="Star History Chart" src="https://api.star-history.com/svg?repos=yinnho/aginxbrowser&type=Date" />
</picture>
</a>
## License
Apache-2.0.
TDQS
Scored across 40 tools
Tool families are cleanly separated: stateless fetch/eval/click are explicitly distinguished from their session_* counterparts, and account_*/render_*/session_* names signal intent. The main fuzzy boundaries are session_challenges vs session_verdict (both report anti-bot/risk status) and the three render_* tools, though descriptions usually resolve them.
Most tools follow predictable families—account_<verb>, session_<operation>, render_<format>—and bare stateless verbs (fetch, search, eval, click) are easy to read. Deviations such as flow_run instead of run_flow and noun-style read tools like session_state/session_network/session_cookies keep it from being perfectly uniform.
40 tools is a heavy surface, especially with 26 session_* variants; even though each is a real browser action, the set will tax an agent's selection and context budget. The server's broad scope explains the count, but it still exceeds what most MCP clients handle gracefully.
The surface covers the full browser workflow: read/search/download, session lifecycle and input, login-state persistence, risk-control verdicts, flow replay, and render/export targets. Minor gaps such as no explicit multi-tab/window management or standalone proxy/user-agent configuration are workaroundable.