Skip to main content
Glama

jev-evidence

JEV Research MCP

A self-hosted, open-source MCP server that uses TypeSafe JEV as a semantic evidence-selection layer for coding-agent web research.

Instead of dumping raw search results and full web pages into an expensive downstream coding model, this server filters twice with JEV — once at the search-result level (before fetching) and once at the content-block level (after parsing) — and returns compact evidence blocks with provenance.

1. What this is

A local MCP server exposing one primary tool, research:

query ("how do I bound concurrent fetches with retries?")
  → search → JEV result filter → fetch selected pages → parse to blocks
  → JEV block filter → context restore → dedup → token budget
  → { sources: [{ title, url, domain, evidence: [...] }], metadata }

It selects evidence. It does not rewrite, summarize, or generate content.

Related MCP server: Wagalo

2. Why it exists

Coding agents burn large amounts of expensive context on irrelevant web content: 20 search hits fetched in full, boilerplate included. JEV (a System One decision model) returns calibrated yes/no judgments with one shared state and many questions per request, which makes it cheap to ask "is this result/block useful?" for dozens of candidates in a single call. Filtering before fetching also avoids network work entirely for discarded sources.

3. Architecture

Coding Agent
    │  MCP tool: research(query, max_sources?, max_evidence_tokens?, mode?)
    ▼
SearchProvider (Brave Search API — live data only)
    │  normalized SearchResult{id,title,url,snippet,domain}
    ▼
JEV result filter — ONE state {task, candidates} + N questions (relevant,
useful, authoritative per candidate), batched (JEV_RESULT_BATCH_SIZE)
    │  deterministic policy: keep = relevant AND useful (conservative band)
    ▼
FetchProvider — plain async HTTP, bounded concurrency (MAX_CONCURRENT_FETCHES),
SSRF-checked, redirect-revalidated, size/timeout/content-type gated
    │  one failed page → diagnostic; never fails the whole request
    ▼
Parser — deterministic HTML → structured blocks (heading/paragraph/code/
list/quote/table), noise stripped, code formatting + hierarchy preserved
    │  oversized blocks split only, with original ID + parent context
    ▼
JEV block filter — per-document batches (JEV_BLOCK_BATCH_SIZE) of
relevant/actionable Noul questions
    │  deterministic policy: keep = relevant AND actionable
    ▼
Assembly — heading context restored → exact dedup (source preference on
duplicates, provenance kept) → whole-block token budget (chars/4 estimate)
    ▼
Compact evidence + provenance + metadata → Coding Agent

Key principle: JEV decides, code disposes. JEV answers semantic questions; all HTTP, parsing, batching, thresholds, budgets, and error handling are ordinary deterministic code. JEV output never mutates state directly — every response is validated (IDs, types, ranges, missing/ duplicate/unexpected answers) before the policy layer reads it.

4. How it works

  1. Search (src/jev_research_mcp/search/): Brave provider returns normalized candidates (URL/title/snippet/domain). Malformed rows are rejected structurally, never semantically.

  2. Result filtering (jev/questions.py + policy/results.py): one shared state plus relevant/useful/authoritative Noul questions per candidate in a single JEV call per batch. Policy keeps relevant AND useful; values in the uncertain band [threshold − margin, threshold) are retained (false negatives hurt more than false positives). Authority only reorders survivors.

  3. Fetch (fetch/fetcher.py, security/ssrf.py): only kept URLs are fetched, concurrently under a semaphore. Every URL (including redirect hops) must use http(s) and resolve to globally-routable IPs — localhost, loopback, RFC-1918, link-local, and cloud metadata addresses are blocked. Responses are streamed under MAX_RESPONSE_BYTES, gated by content type (binary rejected), and capped at MAX_DOCUMENT_CHARS.

  4. Parse (parse/parser.py): script/style/nav/cookie-banner/footer noise removed structurally; headings, paragraphs, code (verbatim, language kept), lists, quotes, tables extracted with parent headings.

  5. Block filtering: blocks batched per document into single JEV calls (relevant/actionable per block), same conservative policy.

  6. Assembly (evidence/): selected blocks get their parent heading back (### Enable streaming + block), code fenced with language, exact duplicates collapsed (winner = highest authority/rank, losers kept as also_seen_at provenance), then a round-robin-per-source whole-block budget under MAX_EVIDENCE_TOKENS (chars/4 approximation — documented, not an exact tokenizer count).

5. Installation

Requires Python ≥ 3.10.

git clone <your-fork-url> jev-research-mcp
cd jev-research-mcp
pip install -e ".[dev]"   # runtime + dev tools (or pip install -e . for runtime only)
cp .env.example .env      # then fill in your own keys (never commit .env)

Verify:

ruff check src tests && ruff format --check src tests
mypy src tests
pytest

6. Configuration

All configuration is environment variables (see .env.example). Every value is validated on load; errors name the variable without echoing secrets.

Variable

Default

Meaning

TYPESAFE_API_KEY

— (required for jev mode)

Your TypeSafe key (JEV_API_KEY also accepted)

TYPESAFE_MODEL

jev-latest

JEV model

TYPESAFE_BASE_URL

https://api.typesafe.ai

API root (gateway-compatible)

SEARCH_PROVIDER

brave

brave (live only)

BRAVE_API_KEY

— (required for Brave)

Your Brave Search key (SEARCH_API_KEY fallback)

MAX_SEARCH_RESULTS

20 (≤ 50)

Search candidates

FETCH_TIMEOUT_SECONDS

15

Per-request timeout

MAX_RESPONSE_BYTES

2000000

Streaming cap

MAX_REDIRECTS

5

Redirect hops

MAX_DOCUMENT_CHARS

200000

Document cap (truncates, never OOMs)

MAX_CONCURRENT_FETCHES

5 (≤ 20)

Fetch semaphore

JEV_RESULT_BATCH_SIZE

20

Candidates per JEV call

JEV_BLOCK_BATCH_SIZE

30

Blocks per JEV call

JEV_MAX_STATE_CHARS

80000

State cap

JEV_MAX_BLOCK_CHARS_FOR_STATE

2000

Per-block text sent to JEV (full text kept for evidence)

MAX_BLOCKS_PER_DOCUMENT

300

Parser cap

RESULT_RELEVANT_THRESHOLD / RESULT_USEFUL_THRESHOLD

0.5

Noul keep thresholds

RESULT_UNCERTAIN_MARGIN

0.15

Conservative retain band

BLOCK_RELEVANT_THRESHOLD / BLOCK_ACTIONABLE_THRESHOLD

0.5

Block thresholds

BLOCK_UNCERTAIN_MARGIN

0.15

Conservative retain band

MAX_EVIDENCE_TOKENS

8000

Evidence budget (approx tokens)

MAX_SOURCES

10

Server-side cap on max_sources

ENABLE_CACHE / CACHE_TTL_SECONDS

true / 3600

Process-local page cache

LOG_LEVEL

INFO

DEBUG for verbose

7. MCP configuration

Stdio transport (V1 supports stdio only):

python -m jev_research_mcp

Example client config — see examples/mcp_config.json (replace /path/to/jev-research-mcp with your checkout; keys are yours):

{
  "mcpServers": {
    "jev-research": {
      "command": "python",
      "args": ["-m", "jev_research_mcp"],
      "cwd": "/path/to/jev-research-mcp",
      "env": {
        "TYPESAFE_API_KEY": "YOUR_TYPESAFE_API_KEY",
        "SEARCH_PROVIDER": "brave",
        "BRAVE_API_KEY": "YOUR_BRAVE_SEARCH_API_KEY"
      }
    }
  }
}

Tool schema — research:

Argument

Type

Required

Notes

query

string

yes

Non-empty, ≤ 2000 chars

max_sources

integer

no

Clamped to server MAX_SOURCES

max_evidence_tokens

integer

no

Clamped server-side; whole blocks only

mode

string

no

"jev" (default) or "baseline" (no JEV, for comparison)

Returns { query, sources: [{ title, url, domain, evidence: [{ text, type, language, parent_heading, block_id, url, score, also_seen_at }] }], metadata } or { query, sources: [], metadata: {}, error: { code, message } } on failure. Works with MCP Python SDK v1 (FastMCP) and v2 (MCPServer).

7a. Local dashboard (visualize the context cut)

python -m jev_research_mcp also starts a localhost dashboard (stdout stays clean for MCP; the URL is printed to stderr):

python -m jev_research_mcp                # MCP stdio + dashboard at http://localhost:8765
python -m jev_research_mcp --ui-only      # dashboard only (preview the UI)
python -m jev_research_mcp --no-ui        # MCP stdio only
python -m jev_research_mcp --ui-port 8899 --open-browser

The page shows the 6 pipeline stages, a live query box, and a kept-vs-cut funnel (search → selected → fetched → blocks seen/kept → evidence → tokens). Everything shown is live — there is no demo data. Without TYPESAFE_API_KEY/BRAVE_API_KEY the header reads needs keys and runs return a structured configuration error pointing at the Configuration section below, where you can save keys (persisted to .env, applied immediately). Config: ENABLE_UI, UI_HOST, UI_PORT, UI_OPEN_BROWSER. No new dependencies (stdlib HTTP server + one static HTML file).

The server starts even with missing keys: misconfiguration comes back as a structured tool error, not a startup crash. The dashboard's separate Configuration section lets you set the provider, API keys, and tuning knobs from the browser — it applies immediately and saves to .env. Keys are never displayed back (blank = keep).

8. Search provider configuration

V1 ships one concrete provider: Brave Search API (GET https://api.search.brave.com/res/v1/web/search, X-Subscription-Token header). Set BRAVE_API_KEY (or SEARCH_API_KEY). Your key is sent only to Brave. To add a provider, implement the SearchProvider protocol (search(query, max_results) -> list[SearchResult]) and wire it in server.build_pipeline. (The test suite uses doubles in tests/helpers.py; production code paths never serve fabricated results.)

9. JEV configuration

Set TYPESAFE_API_KEY (SDK default name; JEV_API_KEY accepted as alias). The adapter uses the official typesafe-sdk (TypeSafeClient / AsyncTypeSafeClient, Noul questions, system_one(state, questions)) against POST {TYPESAFE_BASE_URL}/v1/systemone. A tiny isolated HTTP fallback (jev/http_provider.py) speaks the same documented contract and is used only if the SDK path is unavailable. The rest of the app depends only on the DecisionProvider interface (decide_questions(state, questions, meta)), so the SDK can evolve without touching pipeline code.

Question contracts are fixed in jev/questions.py (result: relevant / useful / authoritative; block: relevant / actionable) — explicit interpretations, no vague "is this good?" prompts.

10. Example usage

# Via MCP (coding agent side): call tool "research"
# { "query": "How do I bound concurrent HTTP fetches with retries in Python?",
#   "max_sources": 6 }
#
# Response (shape):
# {
#   "query": "...",
#   "sources": [{
#     "title": "Async Fetch Guide",
#     "url": "https://example.com/guide",
#     "domain": "example.com",
#     "evidence": [{
#       "text": "### Enable streaming\n\nSet this option to true.",
#       "type": "paragraph", "language": None,
#       "parent_heading": "Enable streaming",
#       "block_id": "b1", "url": "https://example.com/guide",
#       "score": 0.91, "also_seen_at": []
#     }]
#   }],
#   "metadata": { "search_results": 20, "selected_results": 6,
#     "documents_fetched": 6, "blocks_seen": 400, "blocks_selected": 50, ... }
# }

Live benchmark (requires both API keys — it measures real evidence selection):

python -m jev_research_mcp.benchmark.bench --out benchmark_outputs/latest.json

11. Privacy model

The repository ships software only. The author's infrastructure receives nothing — there is no telemetry, analytics, update check, or phone-home in the code or its dependencies' usage (network calls go to exactly three places: your configured search provider, the TypeSafe JEV API, and the source URLs returned by search). Queries, documents, and keys stay on your machine except where necessarily sent to those providers:

  • search query + max_results → your search provider (Brave) to retrieve candidates;

  • task + candidate/block text → TypeSafe JEV API for yes/no judgments;

  • HTTP GET → each selected source URL to fetch content.

No other network destinations are contacted. Logs contain counts, timings, and truncated queries — never API keys or authorization headers.

12. Security considerations

  • SSRF: search results are untrusted. Only http(s) URLs without credentials are fetched; hostnames are resolved and every A/AAAA record must be globally routable (ip.is_global covers loopback, private, link-local, reserved, multicast for v4/v6); metadata-service literals are denied; every redirect hop is revalidated under MAX_REDIRECTS.

  • Bounds everywhere: response bytes streamed under a cap, document chars capped, JEV state capped, concurrency semaphored, retries bounded with backoff (transient only — never auth/validation/security errors).

  • Untrusted inputs validated: search JSON, JEV answers (IDs, Noul ranges, missing/duplicate/unexpected), fetched HTML (never executed, parsed as text only), tool arguments (empty/oversized rejected, limits clamped).

  • Secrets: from your environment only; error messages and logs never echo them (covered by tests).

13. Testing

pytest                      # full suite (~160 tests; offline via tests/helpers.py doubles)
pytest tests/test_live_smoke.py   # live JEV check — only runs with TYPESAFE_API_KEY

Coverage: config/bounds, search normalization, batching, SSRF + redirect validation, fetcher (size/timeout/content-type/errors), parser cases (noise vs. docs/code/lists/tables), question batching shape, JEV validation (missing/unexpected/duplicate/out-of-range), deterministic policies (keep/drop/uncertain), dedup/tokens/assembly (context, fences, budget, provenance), pipeline invariants (discarded results never fetched; one failure ≠ total failure), MCP protocol (advertise/schema/execute/structured errors), dashboard (status/research/config, no secret leaks), benchmark math, cache. Live checks (test_live_smoke, benchmark comparison) skip without keys.

14. Benchmarking

python -m jev_research_mcp.benchmark.bench compares baseline (no JEV) vs jev on identical live queries and writes JSON with date, model, config, thresholds, and per-mode measurements (results, fetches, blocks before/after, approx tokens, latency, JEV requests/tokens). It requires TYPESAFE_API_KEY and BRAVE_API_KEY and exits non-zero without them — there is no fixture mode, so every number reflects real evidence selection.

Savings depend on query, sources, and thresholds: run the benchmark for your workload and quote your own numbers.

15. Limitations

  • JEV is a hosted API: jev mode needs a key and network; mode="baseline" still needs a Brave key (it skips only the JEV filtering).

  • V1 does single-pass retrieval only (no adaptive second pass).

  • Dedup is exact-match only (whitespace-normalized); no near-duplicate or semantic similarity.

  • Token counts are a chars/4 approximation, not a real tokenizer.

  • Fetcher is plain HTTP: JavaScript-rendered pages yield little or no content.

  • Brave is the only bundled search provider (interface is open for more).

  • MCP transport is stdio only.

16. Troubleshooting

Symptom

Cause / fix

Missing TYPESAFE_API_KEY

Set it in .env/environment for jev mode; or use mode="baseline"

Brave search rejected the API key (401/403)

Check BRAVE_API_KEY (or SEARCH_API_KEY)

JEV rejected the API key

Check TYPESAFE_API_KEY; test with tests/test_live_smoke.py

All N selected pages failed to fetch

Sources unreachable/blocked; check network, SSRF logs at DEBUG

Empty sources with search_results > 0

Thresholds too strict for the query; lower *_THRESHOLD or widen margin

JEV rate limit exceeded

Bounded retries exhausted; wait and retry

MCP tool returns validation_error for mode

Use "jev" or "baseline" exactly

17. Development setup

pip install -e ".[dev]"
ruff check src tests && ruff format --check src tests
mypy src tests
pytest
python -m build   # or: pip wheel . --no-deps  (validates packaging)

Commit discipline: one subsystem per commit (feat: …, test: …, docs: …), verified with status/diff/tests/lint/type-check before each commit. No secrets in the tree (.env, caches, logs, benchmark_outputs/ are gitignored).

Available Tools

1 tool
researchB

Research a coding question and return compact evidence with provenance.

Args: query: The coding/research question (non-empty, max 2000 chars). max_sources: Max sources to fetch (1..server cap, default server-side). max_evidence_tokens: Evidence budget cap (approx tokens, whole blocks only). mode: 'jev' (semantic filtering) or 'baseline' (no JEV, for comparison).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNojev
queryYes
max_sourcesNo
max_evidence_tokensNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does add real context — mode semantics ('jev' = semantic filtering, 'baseline' = no JEV) and that the token budget truncates to whole blocks — but says nothing about cost, latency, network/auth requirements, or whether it is a read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-line purpose is front-loaded, followed by a tight per-argument list. It is slightly redundant with the schema's parameter names, but every line adds a constraint or default rather than restating a type.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and all four parameters are covered with constraints and defaults. The remaining gap is usage/selection guidance, which is captured in its own dimension rather than blocking correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it supplies the query length limit (max 2000, non-empty), the max_sources range (1..server cap) and server-side default, the token-budget semantics (approx tokens, whole blocks only), and the two valid mode values, none of which appear in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Research a coding question') plus the output shape ('compact evidence with provenance'). No siblings exist to differentiate from, so a 4 is the ceiling here.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative-tool guidance. The only hint is that 'baseline' mode exists 'for comparison', which implies A/B usage but never states when an agent should prefer one mode over the other.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedresearch

TDQS

A3.5/5.0

Scored across 1 tool

Disambiguation5/5

There is only one tool, so there is no possibility of confusing it with another. Its purpose (answer a coding question with sourced evidence) is unambiguous.

Naming Consistency4/5

With a single tool named 'research', there is no pattern to violate, but it is a bare noun rather than a verb_noun convention. Consistency is trivially satisfied.

Tool Count3/5

One tool for an entire research server is borderline thin; a single rich entry point can work, but there is no separate way to list sources, refine results, or fetch a specific document. It sits at the low end of reasonable scoping.

Completeness3/5

The tool covers the core research lifecycle (query, sources, evidence, provenance, modes), but notable gaps remain: no way to retrieve or expand a specific source, no follow-up/refinement operations, and no metadata or health operation. Agents can work around these but with friction.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to perform web research over MCP: search with a real browser, fetch JS-rendered pages, download PDFs and convert them to Markdown, then search and page through large documents using bounded previews so the context window isn't flooded.
    BSD Zero Clause
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to run local-first web research: intent-routed search across independent engines with reranking, a multi-stage fetch/crawl ladder, and document extraction. Results come back as signed-cursor, citation-bearing evidence envelopes, with an optional separately enabled profile for browser click/type actions.
    AGPL 3.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to retrieve verified web evidence via headless Chromium rendering, sitemap navigation, and search candidate discovery, returning clean Markdown, structured JSON, and verbatim quoted proof.
    161 npm
    MIT