Skip to main content
Glama

Gemini + Jev Bridge for Codex

English | Русский

A local MCP server that lets Codex delegate substantial routine drafting, analysis and code proposals to Gemini. TypeSafe Jev checks the draft against explicit criteria. Codex plans the work, reviews and applies changes, runs tests and gives the final answer.

Gemini uses the included quota of a Google AI Pro subscription through Antigravity CLI. The bridge requires useG1Credits=false: when quota runs out, Gemini work stops until it resets. It does not use a Gemini API key, API billing, credit overages or a paid fallback. Model availability and limits depend on the signed-in Google account and Google's current plan terms.

Jev uses a separate TypeSafe account, API key and quota. It is not included in Google AI Pro. This is an independent project, unaffiliated with Google, OpenAI or TypeSafe.

EN/RU guides

Guide

English

Русский

Installation, usage and troubleshooting

User guide

Инструкция пользователя

Setup prompt and instructions for another GPT/Codex chat

GPT guide

Инструкция для GPT

Context selection, caching and metrics

Performance guide

Быстродействие

Preparing and publishing a clean source export

Publishing guide

Публикация

Related MCP server: Antigravity Codex MCP

Quick start

You need Node.js 22+, npm, Codex with local MCP support, Antigravity CLI signed in to your Google AI Pro account, and a TypeSafe key. These instructions use Windows/PowerShell; live service access on macOS/Linux has not been verified for this release.

In the repository folder:

npm.cmd ci --ignore-scripts
node setup.mjs

Edit the generated settings.json: select an available Gemini Flash model from agy models, set the CLI path and allow only the project folders you need in sourceRoots. Store the TypeSafe key separately in .secrets/TYPESAFE_API_KEY. Then:

npm.cmd test
node smoke.mjs --local-only
node install.mjs

Review the installation plan before applying it:

node install.mjs --apply

The installer registers the MCP server, adds delegation rules and restricts the global Antigravity CLI profile, including manual CLI sessions. It backs up existing files. See the user guide for the exact changes and removal steps, then restart Codex.

MCP tools

Tool

Purpose

delegate_task

Gemini drafts text or a patch; Jev checks 1–6 criteria; results are saved locally

jev_evaluate

Up to 12 compact routing, ranking or evidence questions per request

bridge_status

Local configuration status without contacting the services

Only task text and explicitly selected files are sent to the providers. Credentials, local settings and generated results stay out of Git. Read the data boundaries and delegation rules before selecting files. Jev scores are advisory; they do not replace review or tests. Delegation still uses Codex for task setup and review, so it does not guarantee a fixed quota saving.

Performance

Select inclusive line ranges with source_selections, give Jev a narrow review_context and mapped review_evidence, and keep the default compact response. Exact successful requests are cached for 10 minutes in memory. Identical concurrent Gemini tasks share one execution; distinct tasks return busy. Full drafts/reviews remain in local artifacts, with timing and actual provider-call metrics. See the performance guide.

Verification

npm.cmd test
npm.cmd run check:public

These checks and node smoke.mjs --local-only do not call Gemini or Jev. A live node smoke.mjs call uses their quotas and requires completed setup. Passing offline checks does not confirm account access or remaining quota.

License

MIT. Dependencies retain their own licenses.

Available Tools

3 tools
bridge_statusA
Read-only

Check local Gemini/Jev integration readiness without making paid requests or returning credentials.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered; the description adds two useful behavioral facts beyond them – no paid API calls are incurred and no credentials are returned. It still doesn't say what the readiness check actually inspects or how failure is reported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the cost and credential caveats are appended tightly to the core purpose rather than padding it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should describe what a readiness result looks like, but it only states what is *not* returned (credentials). For a simple zero-param diagnostic this is close to adequate, but the agent cannot anticipate the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. The description correctly implies nothing is configurable and no argument is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) and resource (local Gemini/Jev integration readiness), which is clearly distinct from the sibling action tools delegate_task and jev_evaluate. It doesn't explicitly name the siblings, but 'readiness' frames it as a diagnostic rather than a work-performing tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'without making paid requests' gives clear context that this is a zero-cost preflight check, implying it should be used before invoking a paid path like jev_evaluate. No explicit when-not or alternative is named, but the operational context is legible.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delegate_taskA

Send one self-contained routine task to Gemini Flash, then batch-check its result with Jev. Supply selected absolute file paths and explicit acceptance criteria. Returns a short preview, full result file and review verdict. Cannot change source files, browse or execute commands. Uses Google quota and TypeSafe API credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
filesNo
contextNo
review_checksYes

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give the generic flags (readOnly=false, openWorld=true, idempotent=false, destructive=false). The description adds real behavior the annotations cannot convey: it cannot modify source files, cannot browse or execute commands, consumes Google quota and TypeSafe API credits, and returns a preview plus full result file plus verdict.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with what the tool does, then inputs, then returns, then limits. Every sentence carries distinct information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the return shape (preview, result file, verdict) and the capability boundaries. It does not flag that non-idempotent, open-world execution means re-invocation re-spends credits and re-runs the task, a small remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 4 params, so the description carries the burden. It clarifies that 'files' are selected absolute paths and that 'review_checks' are acceptance criteria, but says nothing about 'task' or 'context' — partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and workflow: delegate one routine task to Gemini Flash, then batch-check the result with Jev. This is clearly distinct from bridge_status (status) and jev_evaluate (standalone review), though it never names the siblings to make the distinction explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives usable selection criteria — 'one self-contained routine task', supply 'explicit acceptance criteria' — and rules out a class of use via the stated limits. It stops short of naming the alternative sibling (jev_evaluate) for direct review without delegation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jev_evaluateA
Read-only

Ask Jev 1-12 narrow typed questions for routing, ranking or evidence checks on Codex/Gemini results. Supply context as text/JSON plus optional selected files. Confidence is advisory, not proof or authorization. Repeated identical evaluations are cached in this server process. Uses TypeSafe API credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNo
contextYes
questionsYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and open-world traits. The description adds genuinely useful context beyond them: confidence is advisory rather than proof/authorization, identical evaluations are cached in-process, and the call consumes TypeSafe API credits. These are real operational disclosures, though permissions and any rate limits are not detailed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core purpose and safely within size limits. The advisory-confidence sentence is slightly tangential but carries useful caveat value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a deep nested question schema and no output schema, the description covers inputs reasonably but does not explain the return shape or how the three question types differ. An agent can invoke it but must reverse-engineer the typed-question contract from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load. It clarifies context as text/JSON, files as optional, and adds a 1-12 question-count constraint not enforced in the schema. However, it does not explain the three question types (choice/noul/score) or the role of instructions vs criteria, leaving the deeply nested oneOf structure largely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: asking Jev 1-12 narrow typed questions for routing, ranking, or evidence checks on Codex/Gemini results. This is clearly distinct from a status check or task delegation. It does not explicitly name bridge_status or delegate_task to differentiate itself, keeping it just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage contexts (routing, ranking, evidence checks) and says to supply context plus optional files, but never states when to prefer this over delegate_task or bridge_status, nor any exclusion conditions. Usage is inferable but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedbridge_status
    • First observeddelegate_task
    • First observedjev_evaluate

TDQS

A3.8/5.0

Scored across 3 tools

Disambiguation4/5

The three tools have largely distinct roles: bridge_status is a readiness probe, delegate_task runs a Gemini task with a Jev review, and jev_evaluate queries Jev directly. There is mild overlap since both delegate_task and jev_evaluate involve Jev evaluation, but the descriptions make the boundaries (full delegated workflow vs. narrow standalone questions) reasonably clear.

Naming Consistency3/5

All names use snake_case, which is good, but the structural pattern is mixed: bridge_status is noun_noun, delegate_task is verb_noun, and jev_evaluate is a namespace-prefixed noun_verb. Readable and not chaotic, but not a predictable verb_noun convention throughout.

Tool Count4/5

Three tools is slightly thin but appropriate for a narrow Gemini/Jev bridge whose scope is status checking, task delegation, and evaluation. Each tool earns its place without redundancy.

Completeness4/5

The surface covers the core bridge lifecycle: check readiness, delegate a task with review, and run standalone evaluations. Minor gaps exist (e.g., no way to cancel, poll, or manage in-flight delegated tasks), but agents can work around these for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers