JEV Research MCP
Integrates Brave Search API as the live web search provider, returning normalized search results (URL, title, snippet, domain) that feed the research pipeline's evidence-selection and fetching stages.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@JEV Research MCPresearch how to bound concurrent fetches with retries in Python"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.

JEV Research MCP
A self-hosted, open-source MCP server that uses TypeSafe JEV as a semantic evidence-selection layer for coding-agent web research.
Instead of dumping raw search results and full web pages into an expensive downstream coding model, this server filters twice with JEV — once at the search-result level (before fetching) and once at the content-block level (after parsing) — and returns compact evidence blocks with provenance.
1. What this is
A local MCP server exposing one primary tool, research:
query ("how do I bound concurrent fetches with retries?")
→ search → JEV result filter → fetch selected pages → parse to blocks
→ JEV block filter → context restore → dedup → token budget
→ { sources: [{ title, url, domain, evidence: [...] }], metadata }It selects evidence. It does not rewrite, summarize, or generate content.
Related MCP server: Wagalo
2. Why it exists
Coding agents burn large amounts of expensive context on irrelevant web content: 20 search hits fetched in full, boilerplate included. JEV (a System One decision model) returns calibrated yes/no judgments with one shared state and many questions per request, which makes it cheap to ask "is this result/block useful?" for dozens of candidates in a single call. Filtering before fetching also avoids network work entirely for discarded sources.
3. Architecture
Coding Agent
│ MCP tool: research(query, max_sources?, max_evidence_tokens?, mode?)
▼
SearchProvider (Brave Search API — live data only)
│ normalized SearchResult{id,title,url,snippet,domain}
▼
JEV result filter — ONE state {task, candidates} + N questions (relevant,
useful, authoritative per candidate), batched (JEV_RESULT_BATCH_SIZE)
│ deterministic policy: keep = relevant AND useful (conservative band)
▼
FetchProvider — plain async HTTP, bounded concurrency (MAX_CONCURRENT_FETCHES),
SSRF-checked, redirect-revalidated, size/timeout/content-type gated
│ one failed page → diagnostic; never fails the whole request
▼
Parser — deterministic HTML → structured blocks (heading/paragraph/code/
list/quote/table), noise stripped, code formatting + hierarchy preserved
│ oversized blocks split only, with original ID + parent context
▼
JEV block filter — per-document batches (JEV_BLOCK_BATCH_SIZE) of
relevant/actionable Noul questions
│ deterministic policy: keep = relevant AND actionable
▼
Assembly — heading context restored → exact dedup (source preference on
duplicates, provenance kept) → whole-block token budget (chars/4 estimate)
▼
Compact evidence + provenance + metadata → Coding AgentKey principle: JEV decides, code disposes. JEV answers semantic questions; all HTTP, parsing, batching, thresholds, budgets, and error handling are ordinary deterministic code. JEV output never mutates state directly — every response is validated (IDs, types, ranges, missing/ duplicate/unexpected answers) before the policy layer reads it.
4. How it works
Search (
src/jev_research_mcp/search/): Brave provider returns normalized candidates (URL/title/snippet/domain). Malformed rows are rejected structurally, never semantically.Result filtering (
jev/questions.py+policy/results.py): one shared state plusrelevant/useful/authoritativeNoul questions per candidate in a single JEV call per batch. Policy keepsrelevant AND useful; values in the uncertain band[threshold − margin, threshold)are retained (false negatives hurt more than false positives). Authority only reorders survivors.Fetch (
fetch/fetcher.py,security/ssrf.py): only kept URLs are fetched, concurrently under a semaphore. Every URL (including redirect hops) must use http(s) and resolve to globally-routable IPs — localhost, loopback, RFC-1918, link-local, and cloud metadata addresses are blocked. Responses are streamed underMAX_RESPONSE_BYTES, gated by content type (binary rejected), and capped atMAX_DOCUMENT_CHARS.Parse (
parse/parser.py): script/style/nav/cookie-banner/footer noise removed structurally; headings, paragraphs, code (verbatim, language kept), lists, quotes, tables extracted with parent headings.Block filtering: blocks batched per document into single JEV calls (
relevant/actionableper block), same conservative policy.Assembly (
evidence/): selected blocks get their parent heading back (### Enable streaming+ block), code fenced with language, exact duplicates collapsed (winner = highest authority/rank, losers kept asalso_seen_atprovenance), then a round-robin-per-source whole-block budget underMAX_EVIDENCE_TOKENS(chars/4 approximation — documented, not an exact tokenizer count).
5. Installation
Requires Python ≥ 3.10.
git clone <your-fork-url> jev-research-mcp
cd jev-research-mcp
pip install -e ".[dev]" # runtime + dev tools (or pip install -e . for runtime only)
cp .env.example .env # then fill in your own keys (never commit .env)Verify:
ruff check src tests && ruff format --check src tests
mypy src tests
pytest6. Configuration
All configuration is environment variables (see .env.example). Every
value is validated on load; errors name the variable without echoing secrets.
Variable | Default | Meaning |
| — (required for | Your TypeSafe key ( |
|
| JEV model |
|
| API root (gateway-compatible) |
|
|
|
| — (required for Brave) | Your Brave Search key ( |
|
| Search candidates |
|
| Per-request timeout |
|
| Streaming cap |
|
| Redirect hops |
|
| Document cap (truncates, never OOMs) |
|
| Fetch semaphore |
|
| Candidates per JEV call |
|
| Blocks per JEV call |
|
| State cap |
|
| Per-block text sent to JEV (full text kept for evidence) |
|
| Parser cap |
|
| Noul keep thresholds |
|
| Conservative retain band |
|
| Block thresholds |
|
| Conservative retain band |
|
| Evidence budget (approx tokens) |
|
| Server-side cap on |
|
| Process-local page cache |
|
|
|
7. MCP configuration
Stdio transport (V1 supports stdio only):
python -m jev_research_mcpExample client config — see examples/mcp_config.json
(replace /path/to/jev-research-mcp with your checkout; keys are yours):
{
"mcpServers": {
"jev-research": {
"command": "python",
"args": ["-m", "jev_research_mcp"],
"cwd": "/path/to/jev-research-mcp",
"env": {
"TYPESAFE_API_KEY": "YOUR_TYPESAFE_API_KEY",
"SEARCH_PROVIDER": "brave",
"BRAVE_API_KEY": "YOUR_BRAVE_SEARCH_API_KEY"
}
}
}
}Tool schema — research:
Argument | Type | Required | Notes |
| string | yes | Non-empty, ≤ 2000 chars |
| integer | no | Clamped to server |
| integer | no | Clamped server-side; whole blocks only |
| string | no |
|
Returns { query, sources: [{ title, url, domain, evidence: [{ text, type, language, parent_heading, block_id, url, score, also_seen_at }] }], metadata }
or { query, sources: [], metadata: {}, error: { code, message } } on failure.
Works with MCP Python SDK v1 (FastMCP) and v2 (MCPServer).
7a. Local dashboard (visualize the context cut)
python -m jev_research_mcp also starts a localhost dashboard (stdout stays
clean for MCP; the URL is printed to stderr):
python -m jev_research_mcp # MCP stdio + dashboard at http://localhost:8765
python -m jev_research_mcp --ui-only # dashboard only (preview the UI)
python -m jev_research_mcp --no-ui # MCP stdio only
python -m jev_research_mcp --ui-port 8899 --open-browserThe page shows the 6 pipeline stages, a live query box, and a kept-vs-cut
funnel (search → selected → fetched → blocks seen/kept → evidence → tokens).
Everything shown is live — there is no demo data. Without
TYPESAFE_API_KEY/BRAVE_API_KEY the header reads needs keys and runs
return a structured configuration error pointing at the Configuration
section below, where you can save keys (persisted to .env, applied
immediately).
Config: ENABLE_UI, UI_HOST, UI_PORT, UI_OPEN_BROWSER. No new
dependencies (stdlib HTTP server + one static HTML file).
The server starts even with missing keys: misconfiguration comes back as a
structured tool error, not a startup crash. The dashboard's separate
Configuration section lets you set the provider, API keys, and tuning
knobs from the browser — it applies immediately and saves to .env.
Keys are never displayed back (blank = keep).
8. Search provider configuration
V1 ships one concrete provider: Brave Search API
(GET https://api.search.brave.com/res/v1/web/search, X-Subscription-Token
header). Set BRAVE_API_KEY (or SEARCH_API_KEY). Your key is sent only to
Brave. To add a provider, implement the SearchProvider protocol
(search(query, max_results) -> list[SearchResult]) and wire it in
server.build_pipeline. (The test suite uses doubles in tests/helpers.py;
production code paths never serve fabricated results.)
9. JEV configuration
Set TYPESAFE_API_KEY (SDK default name; JEV_API_KEY accepted as alias).
The adapter uses the official typesafe-sdk (TypeSafeClient /
AsyncTypeSafeClient, Noul questions, system_one(state, questions))
against POST {TYPESAFE_BASE_URL}/v1/systemone. A tiny isolated HTTP fallback
(jev/http_provider.py) speaks the same documented contract and is used only
if the SDK path is unavailable. The rest of the app depends only on the
DecisionProvider interface (decide_questions(state, questions, meta)),
so the SDK can evolve without touching pipeline code.
Question contracts are fixed in jev/questions.py (result: relevant / useful
/ authoritative; block: relevant / actionable) — explicit interpretations, no
vague "is this good?" prompts.
10. Example usage
# Via MCP (coding agent side): call tool "research"
# { "query": "How do I bound concurrent HTTP fetches with retries in Python?",
# "max_sources": 6 }
#
# Response (shape):
# {
# "query": "...",
# "sources": [{
# "title": "Async Fetch Guide",
# "url": "https://example.com/guide",
# "domain": "example.com",
# "evidence": [{
# "text": "### Enable streaming\n\nSet this option to true.",
# "type": "paragraph", "language": None,
# "parent_heading": "Enable streaming",
# "block_id": "b1", "url": "https://example.com/guide",
# "score": 0.91, "also_seen_at": []
# }]
# }],
# "metadata": { "search_results": 20, "selected_results": 6,
# "documents_fetched": 6, "blocks_seen": 400, "blocks_selected": 50, ... }
# }Live benchmark (requires both API keys — it measures real evidence selection):
python -m jev_research_mcp.benchmark.bench --out benchmark_outputs/latest.json11. Privacy model
The repository ships software only. The author's infrastructure receives nothing — there is no telemetry, analytics, update check, or phone-home in the code or its dependencies' usage (network calls go to exactly three places: your configured search provider, the TypeSafe JEV API, and the source URLs returned by search). Queries, documents, and keys stay on your machine except where necessarily sent to those providers:
search query +
max_results→ your search provider (Brave) to retrieve candidates;task + candidate/block text → TypeSafe JEV API for yes/no judgments;
HTTP GET → each selected source URL to fetch content.
No other network destinations are contacted. Logs contain counts, timings, and truncated queries — never API keys or authorization headers.
12. Security considerations
SSRF: search results are untrusted. Only
http(s)URLs without credentials are fetched; hostnames are resolved and every A/AAAA record must be globally routable (ip.is_globalcovers loopback, private, link-local, reserved, multicast for v4/v6); metadata-service literals are denied; every redirect hop is revalidated underMAX_REDIRECTS.Bounds everywhere: response bytes streamed under a cap, document chars capped, JEV state capped, concurrency semaphored, retries bounded with backoff (transient only — never auth/validation/security errors).
Untrusted inputs validated: search JSON, JEV answers (IDs, Noul ranges, missing/duplicate/unexpected), fetched HTML (never executed, parsed as text only), tool arguments (empty/oversized rejected, limits clamped).
Secrets: from your environment only; error messages and logs never echo them (covered by tests).
13. Testing
pytest # full suite (~160 tests; offline via tests/helpers.py doubles)
pytest tests/test_live_smoke.py # live JEV check — only runs with TYPESAFE_API_KEYCoverage: config/bounds, search normalization, batching, SSRF + redirect
validation, fetcher (size/timeout/content-type/errors), parser cases
(noise vs. docs/code/lists/tables), question batching shape, JEV validation
(missing/unexpected/duplicate/out-of-range), deterministic policies
(keep/drop/uncertain), dedup/tokens/assembly
(context, fences, budget, provenance), pipeline
invariants (discarded results never fetched; one failure ≠ total failure),
MCP protocol (advertise/schema/execute/structured errors), dashboard
(status/research/config, no secret leaks), benchmark math, cache.
Live checks (test_live_smoke, benchmark comparison) skip without keys.
14. Benchmarking
python -m jev_research_mcp.benchmark.bench compares baseline (no JEV)
vs jev on identical live queries and writes JSON with date, model, config,
thresholds, and per-mode measurements (results, fetches, blocks
before/after, approx tokens, latency, JEV requests/tokens). It requires
TYPESAFE_API_KEY and BRAVE_API_KEY and exits non-zero without them —
there is no fixture mode, so every number reflects real evidence selection.
Savings depend on query, sources, and thresholds: run the benchmark for your workload and quote your own numbers.
15. Limitations
JEV is a hosted API:
jevmode needs a key and network;mode="baseline"still needs a Brave key (it skips only the JEV filtering).V1 does single-pass retrieval only (no adaptive second pass).
Dedup is exact-match only (whitespace-normalized); no near-duplicate or semantic similarity.
Token counts are a chars/4 approximation, not a real tokenizer.
Fetcher is plain HTTP: JavaScript-rendered pages yield little or no content.
Brave is the only bundled search provider (interface is open for more).
MCP transport is stdio only.
16. Troubleshooting
Symptom | Cause / fix |
| Set it in |
| Check |
| Check |
| Sources unreachable/blocked; check network, SSRF logs at |
Empty | Thresholds too strict for the query; lower |
| Bounded retries exhausted; wait and retry |
MCP tool returns | Use |
17. Development setup
pip install -e ".[dev]"
ruff check src tests && ruff format --check src tests
mypy src tests
pytest
python -m build # or: pip wheel . --no-deps (validates packaging)Commit discipline: one subsystem per commit (feat: …, test: …, docs: …),
verified with status/diff/tests/lint/type-check before each commit. No secrets
in the tree (.env, caches, logs, benchmark_outputs/ are gitignored).
Available Tools
1 toolresearchB
Research a coding question and return compact evidence with provenance.
Args: query: The coding/research question (non-empty, max 2000 chars). max_sources: Max sources to fetch (1..server cap, default server-side). max_evidence_tokens: Evidence budget cap (approx tokens, whole blocks only). mode: 'jev' (semantic filtering) or 'baseline' (no JEV, for comparison).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | jev | |
| query | Yes | ||
| max_sources | No | ||
| max_evidence_tokens | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does add real context — mode semantics ('jev' = semantic filtering, 'baseline' = no JEV) and that the token budget truncates to whole blocks — but says nothing about cost, latency, network/auth requirements, or whether it is a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line purpose is front-loaded, followed by a tight per-argument list. It is slightly redundant with the schema's parameter names, but every line adds a constraint or default rather than restating a type.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and all four parameters are covered with constraints and defaults. The remaining gap is usage/selection guidance, which is captured in its own dimension rather than blocking correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and largely does: it supplies the query length limit (max 2000, non-empty), the max_sources range (1..server cap) and server-side default, the token-budget semantics (approx tokens, whole blocks only), and the two valid mode values, none of which appear in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Research a coding question') plus the output shape ('compact evidence with provenance'). No siblings exist to differentiate from, so a 4 is the ceiling here.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative-tool guidance. The only hint is that 'baseline' mode exists 'for comparison', which implies A/B usage but never states when an agent should prefer one mode over the other.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
research
TDQS
Scored across 1 tool
There is only one tool, so there is no possibility of confusing it with another. Its purpose (answer a coding question with sourced evidence) is unambiguous.
With a single tool named 'research', there is no pattern to violate, but it is a bare noun rather than a verb_noun convention. Consistency is trivially satisfied.
One tool for an entire research server is borderline thin; a single rich entry point can work, but there is no separate way to list sources, refine results, or fetch a specific document. It sits at the low end of reasonable scoping.
The tool covers the core research lifecycle (query, sources, evidence, provenance, modes), but notable gaps remain: no way to retrieve or expand a specific source, no follow-up/refinement operations, and no metadata or health operation. Agents can work around these but with friction.
Related MCP Connectors
MCP-native web evidence and claim verification: cited, source-grounded evidence for AI agents.
Agent-native search engine with live web research optimized for AI agents.
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform web research over MCP: search with a real browser, fetch JS-rendered pages, download PDFs and convert them to Markdown, then search and page through large documents using bounded previews so the context window isn't flooded.BSD Zero Clause
- AlicenseNot gradedqualityAmaintenanceEnables coding agents to run local-first web research: intent-routed search across independent engines with reranking, a multi-stage fetch/crawl ladder, and document extraction. Results come back as signed-cursor, citation-bearing evidence envelopes, with an optional separately enabled profile for browser click/type actions.AGPL 3.0
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to retrieve verified web evidence via headless Chromium rendering, sitemap navigation, and search candidate discovery, returning clean Markdown, structured JSON, and verbatim quoted proof.161 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to run source-bound evidence checks and bounded batch judgments for classification, extraction, and decision tasks via TypeSafe Jev.999 npmMIT