Skip to main content
Glama
derwells

sieve

by derwells

sieve

sieve is a local MCP server that puts Jev, TypeSafe's System One model, in a coding agent's hands. Jev returns probabilities over a closed set of candidates and never generates text. The general tool is jev_ask: the agent writes the question, enumerates the answer space, passes the items, and gets one answer per item. The rest — jev_grep, jev_rank, jev_search, jev_verify, jev_route, jev_triage_* — are shortcuts for recurring shapes over the same core, with the candidate enumeration or the criteria already written. The agent still opens and reads the survivors itself.

The server runs locally but sends file previews and candidate text to the TypeSafe API, so a TypeSafe API key is required. Get one at https://docs.typesafe.ai/introduction/quickstart.

See Cost notes for pricing and Status for the recall gate results.

Install

Requires Python 3.12 or later and uv.

git clone https://github.com/derwells/sieve.git
cd sieve
uv sync

sieve reads the TypeSafe API key from TYPESAFE_API_KEY in its environment. The bin/sieve-mcp launcher also sources an env file if one exists, so the key never has to sit in a harness config file. The default path is ~/.config/sieve/env; SIEVE_ENV_FILE overrides it. The launcher sources this file as shell code and exports its assignments; values in the file override existing environment values.

mkdir -p ~/.config/sieve
cat > ~/.config/sieve/env <<'EOF'
TYPESAFE_API_KEY=...
BRAVE_API_KEY=...
EOF
chmod 600 ~/.config/sieve/env

Replace ... with your API key. BRAVE_API_KEY is optional and only affects jev_search; remove that line if you do not use Brave, since even the placeholder selects the Brave backend. To search through a local SearXNG instance instead, see Local SearXNG. If no env file exists, the launcher uses whatever is already in the environment.

sieve runs two ways. Over stdio, each harness session spawns its own server process through bin/sieve-mcp. Served, one long-running process listens for streamable HTTP on loopback and every session on the machine shares it. With many concurrent agent sessions the shared server saves a process each (roughly 50 MB apiece); stdio needs no service and remains the fallback.

Shared server

bin/sieve-mcp --http sources the env file as above and serves streamable HTTP at http://127.0.0.1:8723/mcp. It binds to 127.0.0.1 only, and the MCP SDK rejects requests whose Host or Origin header is not loopback. --port or SIEVE_HTTP_PORT changes the port. The server is stateless: no tool keeps per-session state, so a client that idles for hours or outlives a server restart keeps working without re-initialising.

Run it as a systemd user unit. contrib/sieve.service is a template; set the launcher path and a PATH that reaches uv (and the claude, codex or paseo binaries if you use the backends that shell out to them):

sed "s|/path/to/sieve|$PWD|" contrib/sieve.service > ~/.config/systemd/user/sieve.service
systemctl --user daemon-reload
systemctl --user enable --now sieve
loginctl enable-linger "$USER"   # keep it running while you are logged out

Then register the URL. Claude Code, user scope:

claude mcp add --scope user --transport http sieve http://127.0.0.1:8723/mcp

Codex, in ~/.codex/config.toml:

[mcp_servers.sieve]
url = "http://127.0.0.1:8723/mcp"

OpenCode, in ~/.config/opencode/opencode.json:

{
  "mcp": {
    "sieve": {
      "type": "remote",
      "url": "http://127.0.0.1:8723/mcp",
      "enabled": true
    }
  }
}

A served process does not share its callers' working directory, so paths must not depend on it: jev_grep needs an absolute path, and jev_verify needs an absolute base_path for relative citation paths. Both return a clear error otherwise. Over stdio, relative paths still resolve against the harness's directory.

Many callers can use one server at once. Each tool call builds its own TypeSafe client and budget, and opens and closes its own connection to the sqlite caches, so no call sees another's state. The answer caches (.sieve-cache/, ~/.cache/sieve/) are keyed on content, so concurrent callers share hits. Repository walks and transcript reads run in worker threads so one large jev_grep does not stall other callers. Stdio processes left over from before the switch share the same cache files safely; sqlite serialises their writes.

Stdio

Register the launcher with a harness. The launcher resolves the repository from its own location, so any clone path works. Replace /path/to/sieve/bin/sieve-mcp with the absolute launcher path in your clone. The launcher requires Bash, readlink -f, and uv on the MCP client's executable search path.

Claude Code, user scope:

claude mcp add --scope user sieve -- /path/to/sieve/bin/sieve-mcp

Codex, in ~/.codex/config.toml:

[mcp_servers.sieve]
command = "/path/to/sieve/bin/sieve-mcp"

OpenCode, in ~/.config/opencode/opencode.json:

{
  "mcp": {
    "sieve": {
      "type": "local",
      "command": ["/path/to/sieve/bin/sieve-mcp"],
      "enabled": true
    }
  }
}

To run the server directly from the clone: uv run sieve (stdio) or uv run sieve --http. Neither sources the env file; export TYPESAFE_API_KEY first or use the launcher.

Related MCP server: jev-mcp

Tools

jev_ask(question, items, kind="judge", yes=None, no=None, options=None, levels=None, top_k=None, threshold=0.0, max_chars=4000, budget_usd=0.50)

One question you wrote, asked of every item you hold. Write the question, enumerate the answer space, call it. kind picks the primitive:

kind

answer space

per item

judge

yes and no criteria

{id, probability}, sorted, cut by top_k and threshold

choose

options as {label: when it applies}

{id, choice, probabilities}

score

levels, an ordered list, lowest first

{id, level, label, expected, probabilities}

Items are plain strings or {id, text}; a missing id falls back to the list position. Each item's text is cut to max_chars, 40 items go in one request, requests run concurrently under the same budget check the other tools use, and every answer is validated client-side before it is returned. Answers are cached in sqlite under ~/.cache/sieve/ask/, keyed on (model, question, item text, criteria), so repeating a question costs nothing.

{
  "results": [{"id": "crash", "probability": 0.98}, {"id": "dark-mode", "probability": 0.03}],
  "kind": "judge",
  "items_scored": 2,
  "tokens": 1032,
  "input_tokens": 1012,
  "output_tokens": 20,
  "requests": 1,
  "cache_hits": 0,
  "cost_usd": 0.000043,
  "budget_exhausted": false
}

Write the criteria concretely and make them mutually exclusive; Jev reads them literally, and vague criteria give probabilities near 0.5 across the board. Put shared context — the traceback, the spec, the goal — in the question rather than repeating it in every item. Add your own none option to choose when no listed label may fit.

jev_grep(question, path, mode="files", top_k=20, threshold=0.5, budget_usd=0.50)

Ranks a repository against a plain-language question. The walk is gitignore-aware and never follows symlinks. mode="files" scores every file using its path and preview. Files longer than 80 lines use their first 40 lines plus a line-numbered outline; shorter files use their full text. Previews have a 6_000-character cap. mode="functions" runs the files pass first, then splits the strongest surviving files into functions with tree-sitter (fixed 60-line chunks where no grammar applies) and scores those.

Returns:

{
  "results": [
    {"path": "sieve/validate.py", "line_start": 1, "line_end": 91, "kind": "file", "probability": 0.91}
  ],
  "units_scored": 137,
  "tokens": 204233,
  "input_tokens": 201411,
  "output_tokens": 2822,
  "requests": 18,
  "cache_hits": 0,
  "cost_usd": 0.008459,
  "budget_exhausted": false
}

Results are sorted by probability. Units below threshold are dropped. Answers are cached in sqlite, keyed on (model, question, unit text, prompt), so a repeated question incurs no Jev charge when all units are cached and those inputs are unchanged. The cache lives in .sieve-cache/ inside the searched repository when sieve can add that pattern to an existing .gitignore (which it modifies if necessary), and under ~/.cache/sieve/<repo-hash>/ otherwise, so it never leaves an untracked directory in someone else's working tree.

jev_rank(question, candidates, top_k=None, threshold=0.0)

Reranks candidates the agent already holds. candidates is [{id, text}]; a missing id falls back to the list position. Returned IDs are strings. top_k=None returns all candidates that meet threshold. Each candidate is cut to 2000 characters, 40 go in one request, and requests run concurrently.

Returns:

{
  "results": [{"id": "postgres", "probability": 0.94}],
  "candidates_scored": 5,
  "tokens": 1204,
  "input_tokens": 1180,
  "output_tokens": 24,
  "requests": 1,
  "cache_hits": 0,
  "cost_usd": 0.00005,
  "budget_exhausted": false
}

jev_route(ask, routes, budget_usd=0.10)

Chooses one route for ask from caller supplied routes, each with id, description, and aliases, or chooses none when no route fits. It asks one Choice over the routes and none. With more than 10 routes, it scores groups of at most 10, keeps the two strongest routes from each group, then asks a final Choice over the survivors and none. Usage reports the number of stages and groups. For more than five groups, the final Choice uses the 10 strongest survivors from the first stage.

Returns:

{
  "choice": "billing",
  "confidence": 0.92,
  "confidence_source": "sdk",
  "probabilities": {"billing": 0.94, "docs": 0.04, "none": 0.02},
  "usage": {"tokens": 410, "input_tokens": 380, "output_tokens": 30, "requests": 1, "cache_hits": 0, "cost_usd": 0.000016, "budget_exhausted": false, "stages": 1, "batches": 1}
}

jev_search(query, top_k=10, variants=3, depth=30)

Searches the web and reranks the results. depth is the candidate pool, the number of deduped hits Jev reranks. It is capped at 50 and is never below top_k. top_k is how many of those hits come back. Each variant asks the backend for an even share of the pool, at least 10 hits. A backend that pages (searxng) asks the original query for the whole pool, because variants overlap heavily and another page costs less than another variant. The pool is interleaved rank by rank across variants, deduped, and cut to depth. A sparse query returns a smaller pool with pool_short: true. sieve never pads the pool and never adds variants to fill it. Code requests 2 to 4 query variants: the original, one with filler words stripped, a reordered rephrase, and docs and github suffixes when the query names software. Duplicate variants are removed, so fewer may run. The variants run concurrently through one backend. Hits are deduped by canonical URL (lowercase host, no www., no default port, no fragment, no tracking parameters), and the survivors are reranked against the original query through the jev_rank path. Only the top k reach the agent. The JSON below is abbreviated: usage also contains input_tokens, output_tokens, and budget_exhausted. Partial backend failures add backend_errors; if every variant fails, the tool raises an error. Backend results are cached in sqlite under ~/.cache/sieve/search/ for SIEVE_SEARCH_CACHE_TTL seconds (default 3600; 0 turns it off). The cache is trimmed to the newest 2,000 entries. It stores queries and snippets, so its directory is kept at mode 700 and the database and any SQLite sidecar files at 600. A cached call shows cached: true and does not count in backend_requests. Jev usage is reported separately in usage.

Returns:

{
  "results": [{"url": "https://docs.typesafe.ai/...", "title": "Re-ranking", "snippet": "...", "probability": 0.93, "engines": ["google"]}],
  "variants": ["typesafe jev rerank cookbook", "..."],
  "backend": "searxng",
  "usage": {"tokens": 12400, "requests": 2, "cost_usd": 0.00052, "cache_hits": 0},
  "depth": 50,
  "hits_requested": [50, 17, 17],
  "hits_found": 99,
  "hits_deduped": 69,
  "candidates_scored": 50,
  "backend_requests": 6,
  "backend_calls": [{"query": "...", "hits": 50, "wall_seconds": 2.0, "usage": {"requests": 3, "pages": 3, "stopped": "count", "unresponsive_engines": []}}],
  "wall_seconds": 3.0
}

String rules produce query variants. The backend supplies each result's url, title and snippet, and for searxng the upstream engines and any published date. Jev assigns a relevance probability to each deduped hit.

jev_verify(records=None, report=None, base_path=None, budget_usd=0.50, support_threshold=0.6, contradict_threshold=0.5)

Checks claims against cited files or web pages. Provide either records or a Markdown report. A record has a claim, optional claim_context, optional kind (fact or recommendation), optional premises as strings, and citations as [{"locator": "...", "quote": "..."}]. A locator is an HTTP URL or a file path. Relative paths resolve against base_path. Recommendations are exempt from a verdict; their premises are checked as facts using the same citations.

For a report, sieve extracts sentences, bullets, and table rows in code. Headings, parent list items, and table headers supply context. These claims carry extraction_uncertain: true, so review their wording before relying on a verdict.

sieve fetches each cited source once, up to 2 MB (25 MB for a PDF), and records the SHA256 hash of its bytes as source_version. A PDF is turned into text in a child process with a 30 s timeout, a 1 GB memory limit and a 200-page cap. The child runs pdftotext (poppler) when it is installed and pypdf otherwise, and the text is NFKC-normalised so ligatures match typed quotes. An encrypted PDF, a scanned PDF with no text layer, or a parse failure is reported as a fetch failure with its reason. Raw PDF bytes are never scored.

It finds supplied quotes by normalised text match (quote characters and dashes folded, whitespace collapsed, no space before closing punctuation, and one closing ., ; or , on the quote ignored unless a digit precedes it), then selects passages near the quote or by lexical overlap. Jev compares each passage with the entire claim and returns probabilities for supports_fully, partially_supports, contradicts, and does_not_address. The result keeps all passage distributions, the passage with the strongest full support, and the largest contradiction probability. A failed fetch is reported as a flag, not as a low support score.

counts groups factual claims by their top verdict. flagged lists claims with fetch, quote, evidence, threshold, or budget flags. The unqualified_factual_relay_blocked field is true if a factual claim has no fetchable citation or any passage reaches contradict_threshold. The default thresholds are provisional until fitted on the verification eval. usage reports tokens, requests, cache hits, cost, and budget exhaustion.

jev_triage_threads(threads, budget_usd=0.50, request_threshold=0.5)

Scores requests for human input across several normalized threads. Each thread has thread_id, chronological events, contract, and optional status and journal_priority. Events have id, ts, role, kind, and text. The contract has title, first_prompt, and human_amendments. Results contain each request, its raw probabilities, source event IDs, coverage, a snapshot pointer, a thread bucket, and usage. The threshold is provisional.

Candidate requests come from assistant prose. For each candidate, separate Noul questions check whether it requests input, whether later human dialogue answers it, whether the assistant withdrew it, and whether the latest assistant statement says it blocks progress. Long dialogue is scored in overlapping windows. With full coverage, no later human turn makes the request unanswered; incomplete coverage makes it unknown. A running agent with newer assistant progress is not marked blocked. If the contract states acceptance criteria, another Noul checks whether exactly one action remains.

jev_triage_paseo(agent_ids, tail=400, budget_usd=0.50, request_threshold=0.5)

Reads local Paseo agent indexes and native Claude or Codex transcripts, then calls the same scorer. When a native transcript is unavailable, it reads paseo logs text and marks coverage as truncated. The tool accepts agent IDs only. It does not accept log text.

sieve-ask: the same path from a shell

bin/sieve-ask is jev_ask for callers that are not MCP clients. Same core, same batching, caching, budget and validation. Items are a JSON list of strings or {id, text} objects on stdin or in --items FILE.

# one yes/no probability per item
printf '%s' '["fix crash on launch","add dark mode"]' | bin/sieve-ask judge \
  "Does the item describe a bug?" \
  --yes "It reports broken or incorrect behaviour." \
  --no "It asks for new behaviour or is not about behaviour."

# one option per item, with the full distribution
bin/sieve-ask choose "Which team owns this ticket?" --items tickets.json \
  --option "web=browser UI, CSS, React" \
  --option "api=HTTP endpoints, auth, database" \
  --option "none=no listed team fits"

# one level on an ordered scale per item, lowest first
bin/sieve-ask score "How badly does this hurt someone using the product today?" --items bugs.json \
  --level "a blemish nobody is blocked by" \
  --level "an annoyance with a workaround" \
  --level "work cannot be completed or the result is wrong"

judge takes --top-k, --threshold; all three take --budget-usd, --max-chars (default 4000 per item), --model, and --no-cache. From Python, sieve.ask.ask(question, items, kind, ...) dispatches on kind, and judge(...), choose(...) and score(...) are async with the same arguments as keywords. bin/sieve-ask sources the same env file as bin/sieve-mcp.

Backends and auto-selection

backend

how

snippets

titles

searxng

a SearXNG instance's JSON API at SIEVE_SEARXNG_URL

real

verbatim

brave

Brave Search HTTP API, needs BRAVE_API_KEY

real

verbatim

claude

headless claude -p, WebSearch results read off the stream-json stream

none

verbatim

codex

headless codex exec, reads the web_search results on its event stream

real (search results)

verbatim

SIEVE_SEARCH_BACKEND picks one. Otherwise sieve uses searxng when SIEVE_SEARXNG_URL is set, brave when BRAVE_API_KEY is set, and claude if neither is set. Only the chosen backend is called; sieve does not fall back to another backend on failure. SIEVE_SEARCH_CMD replaces the CLI command line, SIEVE_SEARCH_CMD_CODEX or SIEVE_SEARCH_CMD_CLAUDE replaces it for one backend only, and SIEVE_SEARCH_TIMEOUT sets the 90 s per-call limit. The CLI backends require an installed, authenticated claude or codex executable.

The CLI backends default to cheap models: codex runs gpt-6-luna at low reasoning effort, read-only and ephemeral (SIEVE_CODEX_SEARCH_MODEL, SIEVE_CODEX_SEARCH_EFFORT), and claude runs claude-haiku-4-5-20251001 (SIEVE_CLAUDE_SEARCH_MODEL). The child uses the CLI's own login. Variables that reroute a CLI to another endpoint, key or model are removed from its environment: every ANTHROPIC_* and CLAUDE_CODE_USE_* variable, plus CLAUDE_CODE_SUBAGENT_MODEL, CLAUDE_CODE_MAX_CONTEXT_TOKENS, API_TIMEOUT_MS, OPENAI_BASE_URL, OPENAI_API_KEY and CLAUDECODE. SIEVE_STRIP_ENV_PREFIXES adds comma-separated prefixes, for example a wrapper's own variables. Without that, a search launched from a session routed to a third-party endpoint would go to that endpoint.

Search chain

sieve.search_chain tries several routes in order, for callers that want a hosted search first and a logged fallback after it. It is separate from jev_search, which still uses one backend.

python -m sieve.search_chain search --routes codex,claude,searxng --count 10 --json "query"
python -m sieve.search_chain preflight --routes codex,claude,searxng --json "query"

From Python, search(query, count, routes, cache_dir=None, env=None) and preflight(query, routes, cache_dir=None, env=None) return the same JSON; asearch and apreflight are the async forms. search returns:

{
  "hits": [{"rank": 1, "title": "...", "url": "https://...", "snippet": "...", "engines": ["google"]}],
  "route": "claude",
  "fallback_from": "codex",
  "errors": {"codex": "codex search exited 1: ..."},
  "degraded": true,
  "observed": true,
  "transcribed": false,
  "config": {"route": "claude", "model": "claude-haiku-4-5-20251001", "effort": null, "cmd_fingerprint": "28ebe4a424ac", "url": null},
  "cache": "hit",
  "seconds": 0.005
}

For each route in order, sieve reads that route's cache, then calls it live, and moves on only if both give nothing. A later route's cached answer is never served while an earlier route works. degraded is true when the answer came from a later route after a hosted route (codex or claude) failed. fallback_from is the first route that failed, and errors gives each failed route's reason. observed is true when the hits came from the provider's own search results: Codex web_search events, Claude WebSearch tool results, or SearXNG's engine results. transcribed is true when they were parsed from the model's reply, which happens only on Codex releases before 0.156.1. When every route fails, hits is empty, route is null, and the command exits 1.

The cache key is the route, its effective configuration (model, effort, command fingerprint, SearXNG URL), the query with case and spacing normalised, and the count. Entries live in ~/.cache/sieve/search-chain/, or --cache-dir, for SIEVE_SEARCH_CACHE_TTL seconds. The chain reads only the per-route command variables, never SIEVE_SEARCH_CMD.

preflight calls every listed route live and concurrently, without reading the cache, and reports {routes: {name: {ok, results, observed, error, seconds, config}}}. A route is ok only if its hits were observed and at least one is an http(s) URL with a title; transcribed links never pass. A passing result is cached, so the next search for the same query and count costs nothing. It does not open pages.

Local SearXNG

SearXNG is a self-hosted metasearch engine. It sends one query to several upstream engines and merges their results. bin/sieve-searxng runs the official container, pinned to ghcr.io/searxng/searxng:2026.9.25-12f8b6515 by digest, and needs Docker. The container is bound to 127.0.0.1 only and has the JSON API enabled:

bin/sieve-searxng start     # detached, --restart unless-stopped
bin/sieve-searxng status    # container state and /healthz
bin/sieve-searxng logs 50
bin/sieve-searxng stop      # removes the container; it stays down after a reboot

The container returns after a reboot as long as Docker itself starts at boot, until you run stop. On first start the script writes ~/.local/share/sieve-searxng/settings.yml and never overwrites it. The settings enable json under search.formats, turn the limiter off (the instance is single-user on loopback), and enable google and yahoo. The secret key goes in secret.env beside it with mode 600. SEARXNG_PORT (default 8888), SEARXNG_HOME, SEARXNG_CONTAINER and SEARXNG_IMAGE override the defaults. To select it, add to ~/.config/sieve/env:

SIEVE_SEARCH_BACKEND=searxng
SIEVE_SEARXNG_URL=http://127.0.0.1:8888

The backend requests up to SIEVE_SEARXNG_MAX_PAGES pages per query (default 3, at most 5). Each request is limited by SIEVE_SEARXNG_TIMEOUT (default 12 s), and all pages for one query share SIEVE_SEARXNG_DEADLINE (default 30 s). Paging stops early once it has the requested hits or a page returns no new URL. Page 1 is retried once on a connection error, timeout or 5xx. A failure on a later page keeps the earlier pages and is recorded in page_errors. Engines that SearXNG reports as unresponsive (CAPTCHA, rate limit, timeout) are listed in each call's usage.unresponsive_engines. If no engine returns anything, the call fails.

Upstream engines rate-limit and CAPTCHA automated traffic, so which engines are healthy changes over time, and an instance needs occasional settings and image updates. The 2026-09-25 trial filled a 50-hit pool on all four test queries. At times, only one or two engines were answering.

CLI backend notes

Since codex-cli 0.156.1, each completed web_search event carries results[] with url, title, snippet and a ref_id. Search results (turnNsearchM) have real snippets. Opened pages (turnNviewM, from open_page or find_in_page) keep their url and title, but their "Total lines: N" placeholder snippet is dropped. Results without a url are skipped. Hits come only from these events. usage reports hits_source: "events" and counts links in the model's own list that no event observed as unobserved_links. Older releases put only the query on the event. For those, and only when no event has results, the backend parses the markdown bullet list the model writes, so titles are transcribed and snippets empty (hits_source: "transcript"). CLI children get an empty stdin, because codex exec reads a non-terminal stdin, which inside the MCP server is the protocol pipe. Claude's usage.server_tool_use.web_search_requests reports 0 even when results come back, so sieve counts tool results instead.

Cost notes

Measured 2026-09-22, same query, three variants each:

backend

wall time

cost

searxng

2.5–5.9 s at depth=50 (measured 2026-09-25)

none beyond Jev, about $0.0005 per call

brave

2.1 s

a fraction of a cent

claude

16.6 s

~$0.15 for the three calls

codex

29.0 s

not measured

The brave and codex rows exclude the backend's own charges: Brave bills separately under its API pricing, and Codex runs under a Codex subscription rather than metered API cost. The claude figure is the API-equivalent cost that Claude Code itself reports for the call, not a separate metered charge.

In this measurement, one headless Claude call used ~65k tokens and cost ~$0.05 with ~13 s wall time, even for a small query. Claude Code's own system prompt and tool schemas are the floor, even with MCP, settings and extra tools stripped off.

sieve calculates Jev costs from input tokens at $0.042 per million, the rate configured in sieve/jev.py. The eval includes a 1,090-file repository at HEAD. A jev_grep files pass over its 1080-file parent snapshot cost $0.054 and 5.6 s cold; over a 190-file repository, $0.0084 and 2.0 s. jev_grep and jev_verify expose a budget_usd cap in the MCP interface. They return partial results with budget_exhausted: true when the budget stops scoring. The check uses estimated token costs, so actual spend can exceed the cap. jev_rank and jev_search expose no budget parameter. Search usage.cost_usd covers Jev reranking; backend charges are separate. These are measured costs, not a current provider price list.

Design rules

  • Jev selects or scores over a closed set. The fixed tools define that set in code; jev_ask lets the caller define it, and enforces the same closure.

  • Batch every question that shares a state into one request.

  • Validate every answer client-side: probabilities cover the offered set and sum to ~1; a Choice (a selection from fixed options) must pick the max-probability option. The API rounds probabilities to two decimals, so a pick up to 0.01 below the reported maximum still counts. Reject answers that fail validation.

  • Relevance floats are filters, not truth. Thresholds are evaluated on real asks, not copied from cookbooks.

  • The API key is read from the environment. It never appears in a harness config file or in this repository.

Out of scope

  • Jev does not generate summaries or write queries. The Codex search backend does use generated text to extract results.

  • Indexing or embeddings. Every call enumerates fresh; the cache covers repeats.

  • Multi-hop code tracing.

Status

Implemented and evaluated:

  • jev_ask served over stdio by uv run sieve, with judge, choose and score exercised live through a real stdio client in tests/test_ask_live.py. Not separately evaluated: the question is the caller's, so its accuracy is the caller's to check.

  • jev_grep, jev_rank, jev_search and jev_verify served over stdio by uv run sieve.

  • The shared server (sieve --http) exercised by 24 concurrent clients in tests/test_http.py, and verified headless with a live jev_ask from Claude Code, Codex and OpenCode.

  • Recall eval on five past asks in four private repositories, written up anonymised in evals/. At the current default of 8 units per request, the files mode gate requires recall@10 at least 0.8 on 4 of 5 asks. Strict recall over every previously existing file edited by the fix reached 3 of 5, so it failed. Relaxed recall over each ask's single primary fix file reached 5 of 5, so it passed. Recall@10 is the fraction of ground truth files retrieved in the top 10. The eval used threshold=0.0; the default filter can omit additional files. The earlier eval led to outline previews and 8 units per request instead of 4. A criteria sweep did not justify another prompt change.

  • Citation eval for jev_verify on 40 hand-built cases from public sources, half true and half altered, in evals/. Thresholds fitted on 20 and tested on the other 20: no altered claim accepted, no true claim flagged, 18 of 20 four-way verdicts correct on each half. Scope alterations come back as contradicts rather than partial support.

  • All three search backends exercised live. On the acceptance query, brave and claude put the right page first; codex missed it and transcribed its links. On 2026-09-25, with codex-cli 0.156.1, the codex backend read 20 observed results for that query from its events and returned 10 hits, all with real snippets.

  • The searxng backend was trialled live on four query shapes, in evals/. Each query filled a pool of 50 deduped candidates with real snippets in at most 7 requests. It also puts the right page first on the acceptance query.

  • Registration verified headless in Claude Code, Codex and OpenCode, and through a Paseo (an agent management app) plugin that injects the server into every agent. That plugin is separate from the installation instructions above.

Roadmap:

  • jev_route(ask): Choice over a code-defined set of routes, for an orchestrator that has to pick a project for an incoming request.

  • A Noul (a probability-valued judgment) gate for auto-approving read-only shell commands that no static rule matches.

  • Sharper jev_grep criteria for large repositories, where 35 files can legitimately answer "would a developer have to open this".

  • jev_triage_threads(events, contract): ranks agent threads for a human briefing. Code enumerates candidate requests for human input; one Noul per request decides whether later human dialogue answered it and whether progress is waiting on it. Shipped with a Paseo adapter; eval in evals/. On 30 private thread snapshots the candidate method did not beat whole-window scoring on the test half (6 of 15 buckets right against 10 of 15), so the chief should treat its output as a filter to inspect, not a ranking to trust.

Stack

Python 3.12, uv, typesafe-sdk (async client), mcp (stdio and streamable HTTP; FastMCP is MCPServer in mcp 2.x), tree-sitter-language-pack, pytest. The brave backend uses httpx2, which the TypeSafe SDK already pins.

Tests: uv run pytest -q -m "not live" for the offline suite. uv run pytest -m live also hits the real API and needs TYPESAFE_API_KEY.

References

  • Docs index: https://docs.typesafe.ai/llms.txt

  • Pipeline constants borrowed from superagents-lab/jev-search: 40 per rerank batch, 0.6 source threshold, 8 results per lane.

  • Answer validation pattern from browser-use/jev-ultrafast.

License

MIT. See LICENSE.

Repository enumeration in sieve/enumerate.py is adapted from keltokhy/jgrep, MIT, Copyright (c) 2026 Khaled Eltokhy. Its license notice is reproduced in NOTICE and in the module itself.

Available Tools

3 tools
jev_grepRank a repository against a questionB

Rank the files, or the functions inside the strongest files, of a repository by how relevant they are to a plain-language question. Returns {path, line_start, line_end, kind, probability} sorted by probability, plus token and cost usage. Enumeration is gitignore-aware; answers are cached per (model, question, unit).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'files' scores whole files; 'functions' then splits the best files.files
pathYesAbsolute path to the repository or directory to search.
top_kNoHow many results to return.
questionYesWhat you are trying to find out, in plain language.
thresholdNoDrop units scoring below this probability.
budget_usdNoStop and return partial results before spending more than this.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It reveals that enumeration is gitignore-aware, results are cached per (model, question, unit), and returns token/cost usage, which is helpful. However, it omits details like side effects (if any), required permissions, or behavior on empty repos, which are not covered elsewhere.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary action and output. The first sentence states the core function and return structure; the second adds two behavioral traits. There is no fluff or repetition, making it efficient and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a rich output schema and 6 parameters. The description covers the main function, output fields, and key behaviors (gitignore, caching). It does not explicitly mention the budget_usd stop behavior or edge cases, but the schema documents those parameters and the output schema covers return details, so the description is sufficiently complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already documented in the schema. The description adds contextual behavior (gitignore-aware path enumeration, caching keyed by question) that relates to parameters but does not add parameter-specific syntax or format details beyond the schema. This meets the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb (rank) and resource (files/functions of a repository) by relevance to a question, and specifies the output format. It does not explicitly differentiate from sibling tools like jev_search or jev_rank, but its function is unambiguous and distinct from typical search tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the siblings jev_search and jev_rank. There is no mention of alternatives, exclusions, or preferred contexts, leaving the agent to infer usage from the name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jev_rankRerank candidates against a questionB

Score each candidate for relevance to a question and return [{id, probability}] sorted by probability, plus token and cost usage. Candidates are truncated to 2000 characters and sent 40 per request.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_kNoReturn at most this many results; null returns all.
questionYesWhat the candidates are being ranked against.
thresholdNoDrop candidates scoring below this probability.
candidatesYesCandidates as [{id, text}]. Missing ids fall back to list position.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful non-obvious constraints: candidates are truncated to 2000 characters, processed 40 per request, and the response includes token/cost usage. These are valuable operational details. It does not mention rate limits, failure modes, or side effects, but for a stateless rerank operation the disclosed behavior is reasonably sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence front-loads the core purpose and output format; the second adds the key processing constraints. Every sentence earns its place, and the structure makes the tool easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool of this complexity, the description covers the essential invocation context: what it returns, how candidates are truncated, and the request batching. The output schema exists to document return values, so that burden is shared. It omits edge cases like maximum candidate count or error handling, but those are not necessary for correct invocation in most agent workflows.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so the baseline is 3. The tool description adds no new meaning beyond the schema: it does not explain top_k or threshold beyond what the schema already says, and the candidate shape is also already documented. The description's output-format reference reinforces the parameter purpose but does not supplement it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Score'), the resource ('each candidate'), and the outcome ('return [{id, probability}] sorted by probability, plus token and cost usage'). This goes beyond the title's 'Rerank' by specifying the exact output shape. However, it does not distinguish the tool from its siblings (jev_search, jev_grep), so an agent cannot tell when ranking should replace searching or grepping.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use jev_rank versus the sibling tools. The description does not mention that this should be used after candidates have already been retrieved, nor does it contrast with jev_search or jev_grep. The intended use is only implied by the phrase 'Score each candidate for relevance to a question', which is not enough for an agent to make a routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedjev_grep
    • First observedjev_rank
    • First observedjev_search

TDQS

A4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool targets a clearly distinct data domain: web search results (jev_search), repository code units (jev_grep), and arbitrary candidate text (jev_rank). Despite all returning probability-ranked output, the input types and descriptions make selection unambiguous.

Naming Consistency5/5

All tools share the jev_ prefix followed by a simple verb: search, grep, rank. This predictable verb-oriented pattern makes the tool names easy to learn and distinguish.

Tool Count5/5

Three tools is minimal but well-scoped for a focused relevance-ranking server. Each tool handles a distinct retrieval or ranking task, and none feels redundant or missing.

Completeness5/5

The set covers the core relevance-ranking workflow: search the web, find relevant code within a repository, and rerank arbitrary candidates. There are no obvious dead ends or critical missing operations for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to discover and evaluate GitHub repositories from natural-language feature descriptions, returning ranked adoption-grade candidates with evidence and quality signals.
    3
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables coding agents to make cheap, fast probabilistic decisions on every turn, with tools for coding-loop checks, review, verification, screening untrusted input, and ranking candidates.
    6
    1,141 npm
    61
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables frontier coding agents to delegate routine probabilistic judgments to TypeSafe Jev, providing calibrated triage signals for failures, attempts, completion, context ranking, findings, risk, and generic evidence-grounded questions.
    7
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables Codex to rank available tools and skills, rerank code and documentation search results, and triage saved output artifacts using Jev-based relevance scoring while leaving final decisions to Codex.
    1
    6
    MIT