Skip to main content
Glama
README.md
# 🧭 scout

[![ci](https://github.com/tools-for-agents/scout/actions/workflows/ci.yml/badge.svg)](https://github.com/tools-for-agents/scout/actions/workflows/ci.yml)

**The agent's web reader.**

Raw HTML is a terrible thing to feed a model β€” a 380 KB Wikipedia page is ~95k tokens of markup for a few KB of prose. `scout` fetches a URL and gives back **clean, readable markdown** (headings, links, code, lists β€” the substance, none of the chrome), typically **~90% smaller** than the HTML. Every page is **cached**, so re-reading is free and your **whole reading history is searchable**.

Part of [`tools-for-agents`](https://github.com/tools-for-agents). **Zero dependencies** β€” Node's built-in `fetch` + a regex "readability-lite" extractor + `node:sqlite` (FTS5). Pairs naturally with [`cortex`](../cortex): scout clips the web, cortex files it into your second brain.

---

## Why

| Without scout | With scout |
|---|---|
| Feed raw HTML to the model β†’ ~95k tokens of `<div>`s | `scout_fetch` β†’ ~5k tokens of clean prose |
| Re-fetch the same page every time you need it | Cached β€” re-reads are free (`--fresh` to bust) |
| "What did that article say about X?" β†’ fetch again, re-read | `scout_search "X"` across everything you've read |
| No memory of what you've researched | A searchable reading history on disk |

## CLI

```bash
scout fetch https://en.wikipedia.org/wiki/Zettelkasten     # β†’ clean markdown (cached)
scout fetch https://example.com/post --tokens 3000         # cap the returned size
scout fetch https://api.example.com/data.json --raw        # skip extraction
scout search "luhmann note linking" -k 5                   # search your reading history
scout links https://news.ycombinator.com --limit 30        # outbound links to crawl next
scout list | scout forget https://old.example.com | scout stats
scout serve                                                # β†’ reading-room web view :7950
```

Cache location: `$SCOUT_DB` (default `./.scout/cache.db`).

## Reading room (`scout serve`)

![scout serve β€” the reading room: a shelf of cached pages and a clean serif reader](docs/web-view.png)

```bash
scout fetch https://en.wikipedia.org/wiki/Zettelkasten     # read a few pages…
scout serve                                                # β†’ http://localhost:7950  (--port to change)
```

A calm, zero-dependency web view of everything scout has read β€” the same cache the agent recalls from:

- **Read a page from here** β€” paste a url into the reading room and scout fetches it, strips it to clean markdown, caches it, and opens it. Until now the web view could only ever show what the CLI or an agent had already fetched β€” it could read your library but never add to it. It writes to the same read-through cache the agent uses, so anything you read here is instantly searchable and instantly recallable by [`recall`](https://github.com/tools-for-agents/recall). Fetching is a `POST` (it reaches the network *and* writes), and only `http(s)` urls are accepted.
- **The shelf** β€” every cached page as a card (title, source, when it was read, `~token` size), newest first.
- **Filter by host** β€” chips above the shelf (each with a count) narrow the list to one site in a click β€” read everything you've kept from `en.wikipedia.org`, or just your own docs β€” and clear back to all.
- **Recent reads** β€” the articles you've opened surface as clickable chips above the shelf (remembered in the browser only, most-recent first); jump back to one in a click, or **clear βœ•** to forget them.
- **Reading overview** β€” click the shelf stats for the whole history in the unit that actually matters: **tokens**. Reading those pages as raw HTML would have cost *~10.1k tokens*; scout kept *~1.5k* β€” so the headline is **~8.7k tokens saved (86% lighter)**, which is the entire pitch, finally shown. Underneath: **where you read** (top hosts, click one to narrow the shelf), **reading over time** (pages per day), and your **heaviest reads** β€” ranked by what the raw HTML *would* have cost (`~3.9k β†’ ~550`), so the pages scout saved you the most on are named. Agents get the same digest from the `scout_overview` MCP tool.
- **Re-read it β€” and find out if it changed** β€” a cache you can't refresh is a fossil: `fresh` has always existed on `fetchUrl` and the reading room never set it, so a page you read a month ago was the page you'd read forever. **↻ re-read** pulls it again, and answers the question that actually matters β€” not *β€œhere it is again”* but ***did it change***: *β€œUnchanged since you read it β€” what you took from it still holds”*, or *β€œThis page **changed** β€” 3 lines added, 1 removed. What you took from it may be out of date.”* Reformatting isn't a change; only the words are. And **βœ• forget** drops a page from the library (arms first, then fires), gone from the shelf *and* from search.
- **Where this page points** β€” every article now shows its outbound links, read straight from the clean markdown scout kept (**no network trip**). Each one says whether you've **already read it** β€” those open from your library β€” and the ones you haven't are **one click from being read**: scout fetches, strips and opens them, and the shelf grows. Following the web from inside your own reading room, an edge at a time.
- **Search** your whole reading history (FTS5 + bm25) with matched terms highlighted.
- **The reader** β€” clean, comfortable long-form: the extracted markdown rendered with real typographic hierarchy (including **images**), in a **paper** or **night** theme.
- **Keep it in cortex** β€” hit **🧠 β†’ cortex** in the reader and the article becomes a note in your [second brain](https://github.com/tools-for-agents/cortex): the clean markdown, cited to the **original page** (an article's source is the web, not scout's copy of it) with a link back to the cached read alongside. This is the `scout fetch | cortex capture` loop the toolkit was designed around β€” read the web, keep what matters β€” finally one click. scout never writes: your browser POSTs to cortex's own `/api/capture` (point it elsewhere with `SCOUT_CORTEX_URL`).
- **Copy markdown** β€” one **⧉ copy markdown** button in the reader lifts the whole article's clean markdown to your clipboard β€” the same tokens an agent would get from `scout_fetch`, ready to paste into a note, a prompt, or [`cortex`](../cortex).
- **More from this site** β€” the foot of every article lists the other pages you've read from the **same host**, newest first, each a click away β€” so following a source you already trust is one hop, not a re-search.
- **Table of contents** β€” any article with a couple of headings gets a **☰ contents** button; open it for an outline of the page, click a heading to jump to that section, and the current section stays highlighted as you scroll.
- **Reading progress** β€” a slim bar across the top of the reader fills as you scroll, so you always know how far through a long piece you are.
- **Keyboard-accessible** β€” every control has a visible focus ring, the article cards open with Tab + Enter (not just the mouse), and icon controls carry aria-labels.
- Read-only and **cache-only** β€” the web view never touches the network; `/api/page` returns 404 for anything not already read.

Try the demo without a network fetch: `node scripts/seed.js` then `scout serve`.

## MCP server (for agents)

```jsonc
{
  "mcpServers": {
    "scout": { "command": "node", "args": ["/abs/path/to/scout/mcp/mcp-server.js"],
               "env": { "SCOUT_DB": "/abs/path/to/.scout/cache.db" } }
  }
}
```

### Tools

| Tool | Use it to… |
|---|---|
| `scout_fetch` | Read a web page as clean, token-budgeted markdown (cached; `fresh` to re-fetch). |
| `scout_search` | Search every page you've already read β€” ranked snippets, no re-fetch. |
| `scout_links` | Extract a page's outbound links (absolute URLs + text) to decide where to go next β€” and say so, loudly, when the page was a 404/403 error page, a binary file, read only up to the cap, or a list `limit` cut short. |
| `scout_list` | Your recent reading history. |
| `scout_reread` | Has a page changed since you read it? Re-fetch and diff against your cached copy. |
| `scout_forget` | Drop a page from the cache. |
| `scout_stats` | Pages cached, bytes stored, last fetch. |

### The research loop (with cortex)

1. `scout_fetch` the page β†’ clean markdown.
2. `scout_search` your history to connect it to what you've already read.
3. `cortex_capture` the useful parts into your second brain, then `cortex_write` distilled, `[[linked]]` notes.
4. Next time, `cortex_search` / `scout_search` recall it instead of fetching the web again.

## How it works

- **Fetch** uses Node's global `fetch` (follows redirects, 20 s timeout, a plain user-agent). The body is **streamed with a size cap** (`SCOUT_MAX_BYTES`, default 5 MB) rather than buffered whole β€” a runaway or hostile page can't spike memory or bloat the cache, and if a page is bigger than the cap scout reads the start and *says so* up front.
- **Extraction** is regex-based readability-lite: strip `<script>/<style>/<nav>/<footer>/…`, pick the densest `<article>`/`<main>`/`<body>` region, convert headings, links (resolved to absolute), **content images** (`<img>`/`<figure>` β†’ markdown, tracking pixels dropped), code, lists, bold/italic, and decode HTML entities. Not a full DOM parse β€” but it reliably turns an article into readable prose at a fraction of the tokens.
- **Cache** is a `node:sqlite` table keyed by URL; the same URL returns instantly unless `fresh`. An FTS5 mirror makes the whole history searchable by **bm25**, filled to a token budget (β‰ˆ4 chars/token) β€” the same discipline as [`lens`](../lens) and [`cortex`](../cortex).
- **An error page is not a page.** A 4xx/5xx still has a body β€” usually a friendly "Oops, try these instead" β€” and it converts to tidy markdown like anything else. Both network surfaces refuse to let it pass for the real thing: `scout_fetch` leads its markdown with the status, and `scout_links` leads its result with an `error` saying the links below are the *error page's* navigation, not the page's (and `scout links` exits **1**, where `scout fetch` exits 0 β€” fetch's job, handing back a labelled body, was still done; "where does this page point" was not).
- **A cut list is not a complete one.** A links result also says the shape of its own answer, the same way `scout_list` does: `count` is how many outbound links the page has, `shown` is how many came back, and `truncated` means the list you are holding was cut β€” by `limit` (100 by default), or because the page itself was read only up to the fetch cap. Both say which, in words, in `note`. So **`0 links`, `100 links` and `100 of 500 links` are three different answers**, and none of them is silent about what it is not. What scout cannot tell you is what the server never said: a *soft* 404 β€” HTTP `200 OK` with a "sorry, not found" body β€” is a page as far as any HTTP client is concerned, here and everywhere else.
- Non-HTML **text** responses (JSON, plain text) are stored verbatim. A **binary** resource (image, PDF, archive, font β€” by content-type, or a body that decoded to mostly replacement bytes) is not: decoding it as text is mojibake, so scout returns a short note saying what it is instead of handing you β€” and its cache and search index β€” garbage.

## The agent toolkit

<!-- toolkit:start β€” generated from tools.json by tools-for-agents.github.io/build/toolkit-block.mjs; edit that, not this -->

`scout` is the **read the web** leg of **[tools-for-agents](https://tools-for-agents.github.io)** β€” an operating system for agents.
Nine zero-dependency, MCP-native tools that form one loop, with a self at its centre:

| | | |
|---|---|---|
| πŸ›°οΈ | [agent-hq](https://github.com/tools-for-agents/agent-hq) | coordinate β€” The company's work, made visible. |
| πŸ”Ž | [lens](https://github.com/tools-for-agents/lens) | read code β€” Read code without reading files. |
| βš’ | [anvil](https://github.com/tools-for-agents/anvil) | run safely β€” Run it before you claim it works. |
| πŸ” | [keep](https://github.com/tools-for-agents/keep) | hold secrets β€” Use a secret without holding it. |
| 🧠 | [cortex](https://github.com/tools-for-agents/cortex) | remember β€” A second brain that outlives the context window. |
| 🧭 | **scout** | **read the web** β€” The web, ~90% lighter. |
| πŸ”» | [prism](https://github.com/tools-for-agents/prism) | read data β€” Read data without reading the blob. |
| β—Ž | [recall](https://github.com/tools-for-agents/recall) | recall it all β€” One query. Every store you have. |
| πŸ‘ | [iris](https://github.com/tools-for-agents/iris) | see β€” Look at what you built. |
| πŸ‘» | [ghost](https://github.com/tools-for-agents/ghost) | *the self at the centre* β€” A self that persists across sessions. Not a tool: it is what the agent **is** while it calls these. |

**Reading this as an agent?** [`/llms.txt`](https://tools-for-agents.github.io/llms.txt) is the map, and
[`/tools.json`](https://tools-for-agents.github.io/tools.json) hands you all **79 MCP tools** β€” every name, every
description, every install command β€” in **one fetch**, without cloning anything.

<!-- toolkit:end -->

MIT licensed.