gemini-jev-bridge
You can use this MCP server to delegate routine Gemini drafting/analysis and Jev evaluations while Codex remains responsible for planning, review, and applying changes.
Check local Gemini/Jev integration readiness with
bridge_statuswithout paid requests or credential exposure.Delegate one self-contained routine task to Gemini Flash, then have Jev check it against 1–6 explicit criteria; returns a short preview, full result file, and review verdict.
Provide selected absolute file paths and context for delegation; only task text and explicitly selected files are sent to providers.
Ask Jev up to 12 narrow typed questions (choice, boolean-like, score) for routing, ranking, or evidence checks, with optional files and cached repeated evaluations.
Use it for drafting, analysis, code proposals, and advisory Jev scoring; it cannot change source files, browse, or execute commands.
Results are saved locally, exact successful requests are cached in memory for 10 minutes, and live calls consume Google/TypeSafe quotas.
Allows delegating routine tasks to Gemini through a Google subscription in Antigravity CLI, with additional credits disabled.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gemini-jev-bridgedelegate a SQL migration draft to Gemini, then validate with Jev"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini + Jev Bridge for Codex
A local MCP server that lets Codex delegate substantial routine drafting, analysis and code proposals to Gemini. TypeSafe Jev checks the draft against explicit criteria. Codex plans the work, reviews and applies changes, runs tests and gives the final answer.
Gemini uses the included quota of a Google AI Pro subscription through Antigravity CLI. The bridge requires useG1Credits=false: when quota runs out, Gemini work stops until it resets. It does not use a Gemini API key, API billing, credit overages or a paid fallback. Model availability and limits depend on the signed-in Google account and Google's current plan terms.
Jev uses a separate TypeSafe account, API key and quota. It is not included in Google AI Pro. This is an independent project, unaffiliated with Google, OpenAI or TypeSafe.
EN/RU guides
Guide | English | Русский |
Installation, usage and troubleshooting | ||
Setup prompt and instructions for another GPT/Codex chat | ||
Context selection, caching and metrics | ||
Preparing and publishing a clean source export |
Related MCP server: Antigravity Codex MCP
Quick start
You need Node.js 22+, npm, Codex with local MCP support, Antigravity CLI signed in to your Google AI Pro account, and a TypeSafe key. These instructions use Windows/PowerShell; live service access on macOS/Linux has not been verified for this release.
In the repository folder:
npm.cmd ci --ignore-scripts
node setup.mjsEdit the generated settings.json: select an available Gemini Flash model from agy models, set the CLI path and allow only the project folders you need in sourceRoots. Store the TypeSafe key separately in .secrets/TYPESAFE_API_KEY. Then:
npm.cmd test
node smoke.mjs --local-only
node install.mjsReview the installation plan before applying it:
node install.mjs --applyThe installer registers the MCP server, adds delegation rules and restricts the global Antigravity CLI profile, including manual CLI sessions. It backs up existing files. See the user guide for the exact changes and removal steps, then restart Codex.
MCP tools
Tool | Purpose |
| Gemini drafts text or a patch; Jev checks 1–6 criteria; results are saved locally |
| Up to 12 compact routing, ranking or evidence questions per request |
| Local configuration status without contacting the services |
Only task text and explicitly selected files are sent to the providers. Credentials, local settings and generated results stay out of Git. Read the data boundaries and delegation rules before selecting files. Jev scores are advisory; they do not replace review or tests. Delegation still uses Codex for task setup and review, so it does not guarantee a fixed quota saving.
Performance
Select inclusive line ranges with source_selections, give Jev a narrow review_context and mapped review_evidence, and keep the default compact response. Exact successful requests are cached for 10 minutes in memory. Identical concurrent Gemini tasks share one execution; distinct tasks return busy. Full drafts/reviews remain in local artifacts, with timing and actual provider-call metrics. See the performance guide.
Verification
npm.cmd test
npm.cmd run check:publicThese checks and node smoke.mjs --local-only do not call Gemini or Jev. A live node smoke.mjs call uses their quotas and requires completed setup. Passing offline checks does not confirm account access or remaining quota.
License
MIT. Dependencies retain their own licenses.
Available Tools
3 toolsbridge_statusARead-only
Check local Gemini/Jev integration readiness without making paid requests or returning credentials.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered; the description adds two useful behavioral facts beyond them – no paid API calls are incurred and no credentials are returned. It still doesn't say what the readiness check actually inspects or how failure is reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the cost and credential caveats are appended tightly to the core purpose rather than padding it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should describe what a readiness result looks like, but it only states what is *not* returned (credentials). For a simple zero-param diagnostic this is close to adequate, but the agent cannot anticipate the response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. The description correctly implies nothing is configurable and no argument is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource (local Gemini/Jev integration readiness), which is clearly distinct from the sibling action tools delegate_task and jev_evaluate. It doesn't explicitly name the siblings, but 'readiness' frames it as a diagnostic rather than a work-performing tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'without making paid requests' gives clear context that this is a zero-cost preflight check, implying it should be used before invoking a paid path like jev_evaluate. No explicit when-not or alternative is named, but the operational context is legible.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegate_taskA
Send one self-contained routine task to Gemini Flash, then batch-check its result with Jev. Supply selected absolute file paths and explicit acceptance criteria. Returns a short preview, full result file and review verdict. Cannot change source files, browse or execute commands. Uses Google quota and TypeSafe API credits.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| files | No | ||
| context | No | ||
| review_checks | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give the generic flags (readOnly=false, openWorld=true, idempotent=false, destructive=false). The description adds real behavior the annotations cannot convey: it cannot modify source files, cannot browse or execute commands, consumes Google quota and TypeSafe API credits, and returns a preview plus full result file plus verdict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with what the tool does, then inputs, then returns, then limits. Every sentence carries distinct information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully names the return shape (preview, result file, verdict) and the capability boundaries. It does not flag that non-idempotent, open-world execution means re-invocation re-spends credits and re-runs the task, a small remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 4 params, so the description carries the burden. It clarifies that 'files' are selected absolute paths and that 'review_checks' are acceptance criteria, but says nothing about 'task' or 'context' — partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and workflow: delegate one routine task to Gemini Flash, then batch-check the result with Jev. This is clearly distinct from bridge_status (status) and jev_evaluate (standalone review), though it never names the siblings to make the distinction explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives usable selection criteria — 'one self-contained routine task', supply 'explicit acceptance criteria' — and rules out a class of use via the stated limits. It stops short of naming the alternative sibling (jev_evaluate) for direct review without delegation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_evaluateARead-only
Ask Jev 1-12 narrow typed questions for routing, ranking or evidence checks on Codex/Gemini results. Supply context as text/JSON plus optional selected files. Confidence is advisory, not proof or authorization. Repeated identical evaluations are cached in this server process. Uses TypeSafe API credits.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| context | Yes | ||
| questions | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and open-world traits. The description adds genuinely useful context beyond them: confidence is advisory rather than proof/authorization, identical evaluations are cached in-process, and the call consumes TypeSafe API credits. These are real operational disclosures, though permissions and any rate limits are not detailed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core purpose and safely within size limits. The advisory-confidence sentence is slightly tangential but carries useful caveat value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a deep nested question schema and no output schema, the description covers inputs reasonably but does not explain the return shape or how the three question types differ. An agent can invoke it but must reverse-engineer the typed-question contract from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It clarifies context as text/JSON, files as optional, and adds a 1-12 question-count constraint not enforced in the schema. However, it does not explain the three question types (choice/noul/score) or the role of instructions vs criteria, leaving the deeply nested oneOf structure largely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: asking Jev 1-12 narrow typed questions for routing, ranking, or evidence checks on Codex/Gemini results. This is clearly distinct from a status check or task delegation. It does not explicitly name bridge_status or delegate_task to differentiate itself, keeping it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage contexts (routing, ranking, evidence checks) and says to supply context plus optional files, but never states when to prefer this over delegate_task or bridge_status, nor any exclusion conditions. Usage is inferable but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
bridge_status - First observed
delegate_task - First observed
jev_evaluate
TDQS
Scored across 3 tools
The three tools have largely distinct roles: bridge_status is a readiness probe, delegate_task runs a Gemini task with a Jev review, and jev_evaluate queries Jev directly. There is mild overlap since both delegate_task and jev_evaluate involve Jev evaluation, but the descriptions make the boundaries (full delegated workflow vs. narrow standalone questions) reasonably clear.
All names use snake_case, which is good, but the structural pattern is mixed: bridge_status is noun_noun, delegate_task is verb_noun, and jev_evaluate is a namespace-prefixed noun_verb. Readable and not chaotic, but not a predictable verb_noun convention throughout.
Three tools is slightly thin but appropriate for a narrow Gemini/Jev bridge whose scope is status checking, task delegation, and evaluation. Each tool earns its place without redundancy.
The surface covers the core bridge lifecycle: check readiness, delegate a task with review, and run standalone evaluations. Minor gaps exist (e.g., no way to cancel, poll, or manage in-flight delegated tasks), but agents can work around these for the stated purpose.
Maintenance
Related MCP Connectors
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
A paid remote MCP for OpenAI Codex context compressor, built to return verdicts, receipts, usage log
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
No-data MCP handoff for local Claude Code to Codex harness moves. $49 lifetime.
Related MCP Servers
- AlicenseAqualityCmaintenanceCoordinates Codex and Antigravity CLI for a structured software development workflow with workspace authorization, Git-based verification, and acceptance criteria.10MIT
- AlicenseAqualityBmaintenanceA controlled Model Context Protocol bridge that lets OpenAI Codex delegate work to Google Antigravity CLI on demand, with project-scoped opt-in and read-only delegation.151MIT
- AlicenseAqualityAmaintenanceLets Codex delegate tasks to the Google Antigravity CLI (agy) with background job execution, live progress tracking, a visual dashboard, and structured error diagnostics.8MIT
- AlicenseAqualityCmaintenanceEnables OpenAI Codex to delegate tasks to Google Antigravity CLI, running them headlessly and polling for results.4MIT