Moreno.Jev
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Moreno.Jevreview this diff for regressions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Structured code review with Jev
Moreno.Jev exposes TypeSafe Jev through a local Model Context Protocol (MCP) server. It works with Codex and other MCP-capable clients, including Pi through its MCP adapter.
The server provides six tools:
jev_review_codeasks Jev for a structured bug triage, category, severity, and human-review signal.jev_review_diffchecks a proposed change for regressions using the same structured review.jev_check_hypothesisestimates whether a debugging hypothesis fits the supplied code and symptom.jev_review_test_outputclassifies compiler, test, and runtime failure output.jev_triage_bug_reportstructures an issue report into type, impact, reproduction clarity, and follow-up signals.jev_evaluatesends custom typed Choice, Score, and Noul questions.
Jev returns structured decisions, not explanations or code patches. The connected coding agent should verify the results and explain any findings.
Requirements
Python 3.10 or newer
uv, or pip and Python's built-in virtual-environment support
A TypeSafe API key available to the MCP process as
TYPESAFE_API_KEY
Related MCP server: Jev MCP
Install
Clone the repository, then install the locked dependencies:
git clone https://github.com/MorenoLand/Moreno.Jev.git
cd Moreno.Jev
uv sync --lockedThe project creates its virtual environment in .venv. To use pip instead, create a virtual environment with python -m venv .venv, activate it for your platform, and run python -m pip install -e ..
Set TYPESAFE_API_KEY in the environment used to launch Codex, Pi, or another MCP client. Do not put the key in this repository or in the MCP configuration. Restart the client after changing a persistent environment variable.
Codex
Add a user-level entry to ~/.codex/config.toml. Replace the paths with the absolute paths to this clone.
Windows:
[mcp_servers.moreno-jev]
command = "C:/path/to/Moreno.Jev/.venv/Scripts/python.exe"
args = ["-u", "-m", "jev_mcp"]
cwd = "C:/path/to/Moreno.Jev"
env_vars = ["TYPESAFE_API_KEY"]macOS or Linux:
[mcp_servers.moreno-jev]
command = "/path/to/Moreno.Jev/.venv/bin/python"
args = ["-u", "-m", "jev_mcp"]
cwd = "/path/to/Moreno.Jev"
env_vars = ["TYPESAFE_API_KEY"]The env_vars entry forwards the named variable from Codex's environment; it does not store the key in TOML. Restart Codex after editing the configuration.
To make the skill available across repositories, copy .agents/skills/jev-review from this checkout into ~/.agents/skills/jev-review (Windows: %USERPROFILE%\.agents\skills\jev-review). Codex and Pi both discover Agent Skills from this user location; Pi can also use the checked-in project skill and invoke it with /skill:jev-review.
Suggested AGENTS.md instruction:
When I ask for Jev review or invoke the Jev skill, use the moreno-jev MCP tools on only the relevant code, diff, debugging hypothesis, test output, or issue report. Report Jev's structured answers and probabilities, then independently verify flagged issues against the repository and give me evidence and a concrete next step. Never include credentials or unrelated private data in the submitted state.Then invoke it with $jev-review or ask for a Jev review.
Pi
Pi uses MCP through the pi-mcp-adapter package. Install it with pi install npm:pi-mcp-adapter, then restart Pi. Add the following to the project .mcp.json or user-global ~/.config/mcp/mcp.json, replacing the paths for your OS:
{
"mcpServers": {
"moreno-jev": {
"command": "/path/to/Moreno.Jev/.venv/bin/python",
"args": ["-u", "-m", "jev_mcp"],
"cwd": "/path/to/Moreno.Jev",
"inheritEnv": false,
"env": {
"TYPESAFE_API_KEY": "${TYPESAFE_API_KEY}"
}
}
}
}On Windows, use a command such as C:/path/to/Moreno.Jev/.venv/Scripts/python.exe. The inheritEnv and env fields use pi-mcp-adapter's environment handling so the server receives only the explicitly forwarded key.
Pi also discovers this skill from the same .agents/skills path. The adapter exposes MCP through its mcp gateway. If Jev tools are not direct tools in your Pi session, search for jev_review_test_output or another task-specific tool with mcp({ search: "jev_review_test_output" }), then call the returned tool name. Use /skill:jev-review to load the workflow.
Other MCP clients
Configure a stdio server with the virtual-environment Python executable, arguments ["-u", "-m", "jev_mcp"], and the repository as its working directory. Forward TYPESAFE_API_KEY through that client's environment mechanism; do not copy the key into shared configuration.
Data handling
The server does not read repository files, persist prompts, or log request bodies or credentials. A tool call sends its supplied state and questions to TypeSafe's Jev API. Use the skill to keep submissions limited to the material relevant to the review.
Available Tools
6 toolsjev_check_hypothesisC
Estimate whether a debugging hypothesis fits the supplied symptom and code; inputs are sent to TypeSafe.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| context | No | ||
| symptom | No | ||
| language | No | unspecified | |
| hypothesis | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does add one genuinely useful behavioral disclosure beyond the schema: inputs are sent to TypeSafe, which signals external data egress the agent could not otherwise know. It says nothing about whether this is a read-only analysis, what side effects occur, or any latency/rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core purpose front-loaded. The TypeSafe clause is a semantically distinct fact tacked on with a semicolon rather than its own line, but nothing is wasted or repetitive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but the tool takes five parameters with 0% schema coverage, has no annotations, and no usage or distinction guidance. Too much of the calling contract is left undocumented for a tool of this shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are five parameters. The description only gestures at three of them (hypothesis, symptom, code) with no format, syntax or expectation detail, and it ignores 'context' and 'language' entirely. With the schema providing no field documentation, the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: estimating whether a debugging hypothesis fits a symptom plus code. That is far more precise than a generic 'review' or 'evaluate' verb. However, it never distinguishes itself from siblings like jev_triage_bug_report or jev_evaluate, so an agent must infer which of these overlapping analysis tools applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use statement, no prerequisites, and no named alternative among the five sibling tools. The phrase 'supplied symptom and code' hints at the input situation but gives no routing guidance for an agent choosing between this and jev_triage_bug_report.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_evaluateC
Evaluate text or structured state with Jev typed questions; submitted state is sent to TypeSafe.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | ||
| questions | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one useful trait — submitted state is sent to TypeSafe (a data-flow/privacy note) — but says nothing about permissions, failure modes, or how questions are processed. A single partial disclosure is not enough for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding. The semicolon cleanly separates the action from the data-flow note. Efficient, though terse relative to the complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but the two required parameters are undocumented at 0% coverage, the nested questions object is opaque, and there are no annotations. For a tool with nested structured input, this description is not complete enough to invoke confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and neither required parameter ('state', 'questions') is documented in the schema. The description adds only marginal meaning: state may be 'text or structured' and questions are 'typed'. The nested questions object with additionalProperties is left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb (evaluate) and a resource (text or structured state), and mentions 'Jev typed questions'. However, it does not distinguish itself from siblings like jev_check_hypothesis or jev_review_test_output; the verb 'evaluate' overlaps heavily with 'review' and 'check', leaving the agent unsure which to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives named. The agent is left to infer how this differs from the five sibling review/check tools, which is risky given the overlapping semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_review_codeC
Run a structured Jev triage on the supplied code; code and context are sent to TypeSafe.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| focus | No | ||
| context | No | ||
| language | No | unspecified |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one genuinely important trait: code and context leave the environment and are sent to TypeSafe (external data transmission). That is real value. It still omits auth requirements, rate limits, and any notion of whether the call is read-only or mutating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core action front-loaded and the data-disclosure caveat appended after the semicolon. Efficient, though it substitutes jargon for useful detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but with no annotations and 0% parameter coverage the description leaves the agent guessing about focus, language, and expected behavior. For a four-parameter tool it does too little.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across four parameters, so the description must compensate. It only obliquely touches "code" and "context"; the roles of `focus` and `language` are never explained, leaving two of four parameters completely undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb ("Run") and a resource ("code"), so the basic action is recoverable. However, "structured Jev triage" is opaque jargon that doesn't tell an agent what is actually produced or how it differs from siblings like jev_review_diff or jev_triage_bug_report. It is adequate but not self-explanatory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the five siblings (jev_review_diff, jev_evaluate, jev_check_hypothesis, etc.). No prerequisites, no exclusions, no routing condition — the agent must infer context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_review_diffB
Check a proposed diff for concrete regressions; the diff and context are sent to TypeSafe.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | Yes | ||
| focus | No | ||
| context | No | ||
| language | No | unspecified |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses that the diff and context are transmitted to an external service (TypeSafe), which is a meaningful data-handling detail. However, it does not state whether the operation is read-only, whether it has side effects, or what kind of analysis is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the core action front-loaded and no padding. It is tight, though its brevity contributes to the missing parameter and usage detail elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the external-service disclosure is a plus. But with 0% schema coverage, the undocumented 'focus' and 'language' parameters leave the definition short of what an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate. It only accounts for 'diff' and 'context' as data sent to TypeSafe, leaving 'focus' and 'language' entirely unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Check a proposed diff') plus the goal ('for concrete regressions'). The word 'diff' hints at the distinction from sibling jev_review_code, but no sibling is named or contrasted explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named. With siblings like jev_review_code and jev_review_test_output, an agent must guess which review tool applies to a diff versus code versus test output.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_review_test_outputC
Classify supplied build or test output; the log and context are sent to TypeSafe.
| Name | Required | Description | Default |
|---|---|---|---|
| output | Yes | ||
| command | No | ||
| context | No | ||
| language | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a non-obvious behavioral trait the schema cannot express — that the log and context are transmitted to TypeSafe (an external data-sharing/privacy fact). However, it omits classification categories, failure behavior, and any limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the action front-loaded and no wasted words. It is terse to the point of being slightly cryptic ('TypeSafe' is undefined), which keeps it just short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but with four parameters at 0% schema coverage, an external data-transfer side effect, and zero annotations, the description leaves the agent without enough to invoke it confidently or understand what it may send off-box.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all four parameters, so the description must compensate and does not. It never explains what 'command' or 'language' should contain, nor the expected form of 'output' or 'context', leaving half the inputs semantically opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Classify supplied build or test output'. An agent can tell this is a classification tool rather than a code/diff reviewer, but the description never explicitly contrasts it with siblings like jev_evaluate or jev_triage_bug_report, so differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives named despite five closely related siblings (jev_review_code, jev_review_diff, jev_evaluate, jev_check_hypothesis, jev_triage_bug_report). The agent gets no routing signal beyond guessing from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_triage_bug_reportC
Triage a supplied bug report into typed category, impact, and reproduction signals.
| Name | Required | Description | Default |
|---|---|---|---|
| report | Yes | ||
| context | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden, yet it only names the output dimensions. It does not state whether the operation is read-only, what permissions are needed, whether it is deterministic, or any rate/cost characteristics. The output schema covers return shape, but the behavioral profile is essentially undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the verb and resource come first and the outputs follow. It is efficient, though its brevity is partly under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but with zero annotation coverage and zero parameter documentation the definition is not complete enough for an agent to call the tool confidently. It omits usage context and the meaning of the optional 'context' input.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and does not. 'Supplied bug report' loosely implies the 'report' parameter, but 'context' (an optional free-text field) is never explained — what it should contain, how it influences triage, or whether it is required. This leaves one of two parameters entirely opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Triage') and resource ('bug report') and enumerates the outputs ('typed category, impact, and reproduction signals'), so the agent knows precisely what the tool produces. It is distinguishable from siblings like jev_review_code or jev_check_hypothesis, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the five sibling tools, no prerequisites, and no exclusions. An agent must guess whether triage should precede or replace jev_evaluate or jev_check_hypothesis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
jev_check_hypothesis - First observed
jev_evaluate - First observed
jev_review_code - First observed
jev_review_diff - First observed
jev_review_test_output - First observed
jev_triage_bug_report
TDQS
Scored across 6 tools
Most tools are distinguished by their input type: code, diff, test output, bug report, hypothesis, or generic state. However, jev_evaluate is broad and could overlap conceptually with any of the specialized review tools, creating some ambiguity.
All tools use a consistent jev_ prefix and snake_case, with most following a verb_noun pattern (review_code, review_diff, check_hypothesis). jev_evaluate is a minor deviation because it lacks a noun target.
Six tools is well-scoped for a focused code-review and triage assistant. Each tool covers a distinct input type and earns its place without bloat.
The surface covers common triage scenarios: code, diffs, test logs, bug reports, hypotheses, and generic evaluation. Minor gaps exist for specialized reviews like security or dependency analysis, but the generic evaluate tool can partially compensate.
Maintenance
Related MCP Connectors
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
A paid remote MCP for AI agent browser DevTools MCP, built to return verdicts, receipts, usage logs,
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
Related MCP Servers
- AlicenseAqualityBmaintenanceProvides coding agents and CI with a typed decision layer that sends bounded state and questions to Jev, then returns deterministic actions for review, risk assessment, requirement checks, and verification.91,141 npmMIT
- AlicenseAqualityBmaintenanceEnables frontier coding agents to delegate routine probabilistic judgments to TypeSafe Jev, providing calibrated triage signals for failures, attempts, completion, context ranking, findings, risk, and generic evidence-grounded questions.7MIT
- AlicenseAqualityCmaintenanceEnables coding or reasoning agents to request structured judgments from TypeSafe's Jev model at decision points, including choices, scores, claim verification, and code reviews, with probabilities and confidence returned as data.5173 npmMIT
- AlicenseAqualityCmaintenanceEnables agents to call typed code-review and content-moderation decision tools, returning structured verdicts, probabilities, and confidence-gated actions.2MIT