codearia-sieve
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codearia-sieveParse https://example.com/article into dates, numbers, and text chunks"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
npx codearia-sieve # MCP server for Claude Code, Cursor and any agent
npm i codearia-sieve # or the libraryWhat it does
An agent that needs a web page fetches the whole thing: navigation, cookie banner, footer, ad slots, a megabyte of framework markup. Then a model paid per token digs through the pile for one paragraph.
codearia-sieve does the digging before the model sees anything — and returns the page as state, not prose:
Dates become dates
"Published September 15, 2026" → "2026-09-15".
Read from JSON-LD, meta tags and <time> first; from a byline only when the markup is silent, and never guessed.
Numbers become facts
"$42 per billion tokens" → { value: 42, unit: "USD_per_billion" }.
Works in nine languages; the decimal comma follows the page's language. A number without a unit is not a fact. A year is never a fact.
Text becomes chunks that fit
Each chunk knows its size in tokens and characters, the blocks it was built from, and the #anchor on the page where a decision can be checked.
Everything else — menus, footers, banners, tag rows, "read more" — is removed, and with trace: true you get the list of what was removed and why.
Related MCP server: cleanfetch
Who it is for
People who build agents and have seen the bill. Every fetched page costs tens of thousands of tokens before the agent has read a word of it. Median page in the sample: 53 718 tokens in, 1 106 out.
People who run cheap decision models. Classifiers, rankers, System One models like Jev that judge instead of write. They are nearly free and very fast, and they have hard edges: they cannot count, they read dates as text, and their accuracy drops as irrelevant material fills the context. Every "page to markdown" tool prepares input for a reader. This one prepares input for a judge.
People who need answers they can check. A verdict from scraped text is unprovable unless each piece points back to its source. Here every fact names its block and every chunk carries an anchor.
How it works
Eight steps of ordinary code. No model runs unless you plug one in. The same HTML gives the same JSON, byte for byte.
Fetch —
robots.txtfirst; a refusal is reported, not bypassed. Plain HTTP, honest user agent.Parse — HTML into a DOM with
linkedom. No browser.Dates and ids from the untouched tree — cleaning strips
<head>, bylines and attributes, so both are read before it runs.Clean — Defuddle removes chrome; Readability takes over if it comes back empty. Inline elements get a space first, so
<span>20 Sept</span><span>10 min</span>never becomes202610 min.Blocks — headings, paragraphs, lists, tables, code, quotes, in order, with positional ids. Old pages set with
<br><br>become paragraphs too; table rows keep their column headers.Facts and anchors — ids restored; numbers become facts only beside a unit; ranges keep both ends.
Chunk — greedy, in order, under two budgets at once: 20 000 tokens and 50 000 characters by default.
Assemble —
state,markdown,usage,warnings, and the trace on request.
const r = await sieve({ kind: 'url', url: 'https://docs.typesafe.ai/models' });
r.state.title // "Models"
r.state.facts[0] // { value: 42, unit: "USD_per_billion", label: "price_btok_mtok",
// context: "Price (per Btok / per Mtok) | jev-1.13.0: $42 / $0.042", from: "b3" }
r.state.facts[1] // { value: 0.042, unit: "USD_per_million", … } — paired by position
r.state.chunks[0] // { id: "c1", tokens: 1210, chars: 5357, anchor: "Current models",
// headings: ["Current models", "Pricing", …], blocks: ["b1", …, "b36"], text: "…" }
r.usage // { rawTokens: 127413, stateTokens: 1211,
// visibleChars: 4939, stateChars: 5357, chunks: 1, ms: 1503 }
r.warnings // []
r.markdown // the same article, for a human or a generative modelExpected outcomes never throw. They come back as warnings, each named:
Warning | Meaning |
| the site asks crawlers to stay out; we did not fetch |
| a bot challenge or a refusal (403, 405, 429, "Just a moment…"), with the status |
| a 404 or a 500 that still rendered an error page; not the page you asked for |
| the page marks its article as not free; you got the teaser |
| the article container is empty and a script would fill it |
| a big page that yielded little prose — a front page, a listing |
| one block exceeded the budget and was cut on sentence boundaries |
| the page has more facts than the 500 listed — a long fee schedule, say |
Use it from an agent
You say what you want in plain words. The agent finds the pages, calls sieve_page for each, hands the state to a decision model with a typed question, and writes up the result. Sieve prepares. The judge judges. The agent writes.
{ "mcpServers": { "sieve": { "command": "npx", "args": ["-y", "codearia-sieve"] } } }Listed in the official MCP Registry as io.github.AntonG87/codearia-sieve; clients that read the registry can install it by name.
sieve_page — url or html
Returns typed structuredContent with an output schema: source, state, usage, warnings. In the default summary mode chunks carry sizes, anchors and their headings but no text — the agent sees the outline of what exists without paying for it. mode: "full" and mode: "markdown" when you want everything.
sieve_chunk — url, id
The text of one chunk from the last result for that URL, no refetch. Overview first, then only what is needed — the tool applies its own idea to itself.
Pairs with jev-mcp: chunks are sized to fit its fields, so state goes straight into a typed question.
Checked by a judge
The claim is that a decision model gets better input from Sieve than from raw text. So the output was handed to one. examples/jev.ts drives both MCP servers with the official client — codearia-sieve prepares six pages (API docs, a release note, two Wikipedia articles in two languages, two pricing pages), Jev judges them through jev-mcp. Same run, 21 September 2026:
Question to Jev | Input from Sieve | Result |
| title + head of the first chunk, under the tool's 2 000-char limit | 6 of 6 correct; 5 auto, 1 flagged for review — a page that is both docs and a rate card |
| every fact as a claim, its chunks as evidence | 11 of 11 verified, all auto, confidence 0.86–1.0 |
| first chunk, a date regex, a description | agrees with Sieve where the page states a date; Sieve also reads JSON-LD and |
The first pass of this test did its job the other way round: Jev sent three facts to review and contradicted one. All four traced to Sieve — a table row labelled by its column header instead of its row header, two rates in one header left unpaired, and a Russian bibliographic "256 с." read as seconds. Fixed, tested, rerun: 11 of 11. A judge that can tell you when your parser is wrong is the point of the whole pairing.
TYPESAFE_API_KEY=… node --experimental-strip-types examples/jev.tsUse it as a library
import { sieve } from 'codearia-sieve';
await sieve({ kind: 'url', url }); // fetch it
await sieve({ kind: 'html', html, url }); // already have it; url only for anchors
await sieve(input, {
budget: { maxTokens: 8000, maxChars: 30000 }, // chunk limits
trace: true, // everything discarded, and why
tokenizer: myTokenizer, // o200k by default; swap for your model's
fetcher: myFetcher, // your transport, or a file reader in tests
now: () => fixedDate, // injected clock: identical output on identical input
selector: mySelector, task: 'is this about pricing?', // relevance judge; nothing runs without one
});parseDate, findDates and the vendor limits (JEV, JEV_MCP, DEFAULT_BUDGET) are exported too.
Where it stops
Front pages, listings and product pages have no article to find. You get the headlines and a
thin-contentwarning, not a fake win.Articles rendered by JavaScript come back as
empty-without-jswhen the container is empty. A site that ships a teaser and streams the rest cannot be told apart without a browser; you get the teaser.Text-heavy pages save less. A whole novel saves 22 %, an RFC 84 %: there is no wrapping to remove and the text is kept in full. That is the tool working.
A pricing grid is not a table. A fact knows the block it came from, not the plan column it sits under; Sieve does not guess the pairing. Send the chunk — a pricing page is about a thousand tokens after cleaning — and let the judge read it:
examples/pricing-watch.ts.Tokens are counted with o200k as an approximation. Pages over a megabyte get a sampled count and
usage.rawTokensEstimated: true.
Develop
npm install
npm test # 71 tests, offline, a few seconds
npm run demo -- <url> # the token bill for one page
npm run bench # the 20-page benchmark set
npm run analytics # the 56-page random sample: rows, CSV, summaryDesign notes — vision and architecture — are in docs/.
Available Tools
2 toolssieve_chunkRead one chunkA
Returns the text of a chunk from a page prepared earlier with sieve_page (same url).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Chunk id such as c3. | |
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| text | Yes | |
| anchor | No | |
| tokens | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses a stateful dependency on a prior sieve_page call with the same URL. However, it does not describe failure behavior, whether the page must still be cached, or any other side effects or requirements beyond that dependency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. It front-loads the action and then provides the key precondition in a clear parenthetical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with only two required parameterschers an existing output schema, and the description covers the core precondition. It lacks error-handling guidance, but for a straightforward read operation this is not a critical omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, with id already documented as a chunk ID. The description adds some extra meaning for url by stating it must be the same URL used with sieve_page, which helps an agent select the correct value. This is modest but not fully comprehensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: it returns the text of a chunk from a page. It also distinguishes itself from its sibling sieve_page by noting the page must have been prepared earlier with sieve_page, so an agent can tell the tools apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the proper usage sequence: call sieve_page first, then sieve_chunk with the same URL. It does not explicitly state when not to use this tool or list alternatives, but with only one sibling and a clear dependency, the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sieve_pagePrepare a page for decisionsA
Turns a web page into decision-ready state: dates as ISO fields, numbers with units as facts, text in token-budgeted chunks with anchors back to the page, and the token bill (raw vs state). Start with mode=summary; fetch chunk text with sieve_chunk or mode=full only when needed.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Page to fetch. Either url or html is required. | |
| html | No | Raw HTML to process instead of fetching. Pair with url for anchors. | |
| mode | No | summary: state without chunk text (cheap, default). full: with chunk text. markdown: readable text instead of state. | summary |
| task | No | What you intend to decide; enables relevance selection when a selector is configured. | |
| trace | No | Include what was discarded and where the date came from. | |
| maxChars | No | Chunk budget in characters. | |
| maxTokens | No | Chunk budget in tokens. |
Output Schema
| Name | Required | Description |
|---|---|---|
| state | Yes | |
| trace | No | |
| usage | Yes | |
| source | Yes | |
| markdown | No | |
| warnings | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description is the sole source of behavioral context - and it provides meaningful details: output is a state with ISO dates, unit-aware facts, anchored chunks, and a raw-vs-state token bill. It also signals cost behavior by recommending the cheap summary mode. It does not fully cover failure, authentication, or network behavior, but it goes well beyond a bare mutation/read label.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the purpose first and the mode guidance second. Every clause earns its place: no filler, repeated schema content, or vague preamble.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no annotations, the description covers the core decision (which mode and when to use sieve_chunk) and points to the output form (state, token bill). Parameter details and return values are covered by the 100%-covered schema and output schema; a short note on when markdown mode is preferred would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already documented and the baseline applies. The description adds workflow advice for mode ('only when needed') but does not add new meaning for url, html, task, trace, maxChars, or maxTokens beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb plus resource ('Turns a web page into decision-ready state') and enumerates concrete transformations: dates as ISO fields, numbers with units as facts, text in token-budgeted chunks with anchors, and a token bill. It also references the sibling tool sieve_chunk to clarify that chunk-text retrieval belongs elsewhere, so an agent can distinguish it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Start with mode=summary; fetch chunk text with sieve_chunk or mode=full only when needed' is explicit guidance on the default workflow and when to switch to the sibling or a more expensive mode. This directly answers when to use this tool versus alternatives, with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.2.2- First observed
sieve_chunk - First observed
sieve_page
TDQS
Scored across 2 tools
The two tools have clearly distinct roles: sieve_page processes a web page into chunks and metadata, while sieve_chunk retrieves the text of a specific chunk. There is no overlap or ambiguity—an agent can easily decide which tool to call based on whether it needs to prepare a page or access its content.
Both tools follow the same verb_noun pattern with the 'sieve_' prefix: sieve_page and sieve_chunk. The naming is perfectly consistent, making the toolset predictable and intuitive.
With only 2 tools, the server feels minimal for the stated purpose of turning web pages into decision-ready state. While the core workflow (prepare and fetch chunks) is covered, the narrow surface may not justify a full server; however, it is within the borderline range where it could be acceptable for a very focused utility.
The primary lifecycle is complete: sieve_page prepares a page and sieve_chunk retrieves its content. Minor gaps exist—such as lacking a way to list already-prepared pages or clear cached data—but these are not critical and can be worked around by tracking URLs externally. The inclusion of mode=full also provides an alternative to chunk retrieval, though it overlaps slightly.
Maintenance
Related MCP Connectors
Give agents eyes on any web page: structured context, and changes explained in plain language.
Web scraping for agents. Point it at a URL and it returns the page as clean markdown, JavaScript-rendered pages included. Point it at a site and it maps the URLs or crawls the section you need in the background, a few pages at a time so results fit in the conversation. Search the web and read full pages, extract fields with a JSON schema you define (validated, never invented), read a store's catalogue or a blog's posts from the platform's own feed, and check whether a page has changed. Failed requests cost nothing. The free plan includes 1,500 credits a month.
- mcpOAuthcom.sequentum
Turn the web into structured, reliable, actionable enterprise data for AI Agents
Clean Markdown and AI-readability scoring for any URL. Built for AI agents.
Related MCP Servers
- AlicenseAqualityNot gradedmaintenanceEnables hybrid web search and intelligent content extraction, combining semantic search with documentation-optimized reading that strips noise and returns clean, token-efficient context for AI agents.2MIT
- AlicenseAqualityCmaintenanceEnables AI agents to read web pages reliably, returning clean markdown content, hyperlinks, and metadata without navigation or ad noise.36 npmMIT
- AlicenseAqualityCmaintenanceConverts raw HTML into structured, AI-readable page maps with 97% token reduction, enabling agents to read, click, type, and navigate any web page.91336AGPL 3.0
- AlicenseAqualityCmaintenanceEnables AI agents to extract clean, structured web content (articles, tables, links, visual layouts) optimized for LLM token efficiency, with fast response times and optional JavaScript support.567 npmMIT