Skip to main content
Glama

gemini2api-mcp

MCP wrapper, credential auto-sync, and service patches for gemini2api — a self-hosted service that turns the Gemini web UI into an OpenAI-compatible API using browser cookies.

This repository contains everything that runs on the client side of such a deployment:

Component

What it does

server.mjs

MCP server exposing a gemini_chat tool (text, images, PDF/Office attachments) to any MCP client — ZCode, Claude Code, Codex, OpenAI agents, …

check.mjs + run-scheduled.mjs + ssh_tunnel.py

Credential auto-sync: keeps the server's Google cookies fresh by reading the signed-in Chrome session, comparing SHA-256 fingerprints via the admin API, and submitting at most one idle-window web-form update

launchagent/

macOS LaunchAgent template that runs the checker every 5 minutes

service-patch/

Hardening and bug-fix patches for the upstream gemini2api service (non-root Docker, masked credentials in admin API, idle-reservation cookie updates, atomic persistence + regression test)

calibration/

Research: how Gemini reports pixel coordinates, and the prompt recipe that makes them pixel-accurate (28/28 targets ≤1px)

Required dependencies / upstream projects

  1. xwteam/gemini2api (upstream service, non-commercial license) The server this repo wraps. Deploy it first; it provides the http://host:5918/openai/v1 endpoint, the /admin management API, and the cookie-based account pool. The files under service-patch/ are derivatives of this project and follow its license.

  2. MCP TypeScript SDK (@modelcontextprotocol/server, npm) — MCP protocol implementation used by server.mjs.

  3. zod (npm) — tool input schemas.

  4. paramiko (Python) — loopback SSH tunnel for management traffic (ssh_tunnel.py).

  5. A Chrome automation bridge speaking the {action, args} JSON protocol on 127.0.0.1:10086 (navigate / fill / click / evaluate / cdp Network.getCookies). The reference setup uses the Kimi WebBridge extension + native daemon (~/.kimi-webbridge/bin/kimi-webbridge); any CDP-based bridge with the same command shape works.

  6. Node.js ≥ 20 (built-in fetch, AbortSignal.timeout) and Python ≥ 3.11. The scheduler/notifier part is macOS-only (LaunchAgent + osascript); the MCP wrapper itself is cross-platform.

Related: UI-Venus-MCP — a cross-platform computer-use MCP that can use this endpoint as a pluggable vision provider (see its docs/gemini-coordinate-calibration.md).

Related MCP server: OpenAI-Compatible MCP Gateway

Quick start

git clone https://github.com/q1820926174-cpu/gemini2api-mcp
cd gemini2api-mcp
npm install
cp .env.example .env   # fill in your server URL + keys
npm start              # stdio MCP server

Register in an MCP client (ZCode ~/.zcode/cli/config.json example):

{
  "mcp": { "servers": {
    "gemini2api": {
      "command": "node",
      "args": ["/absolute/path/to/gemini2api-mcp/server.mjs"]
    }
  }}
}

Tool surface (gemini_chat):

{
  "prompt": "What is in this image?",
  "model": "gemini-flash",              // or gemini-pro / -lite / -thinking
  "system": "optional system instruction",
  "images": ["/abs/path/pic.png"],       // up to 8, image types
  "attachments": ["/abs/path/doc.pdf"],  // up to 8, pdf/text/office
  "max_tokens": 4096
}

Limits: single-round only (no conversation state), max 8 files / 20 MiB per call, 180 s upstream timeout (image calls take 5–120 s; MCP client timeouts below ~180 s will cut image requests short).

Credential auto-sync (optional)

Google rotates the __Secure-1PSID* cookies behind gemini2api; when that happens the service starts failing until someone re-pastes cookies into the admin dashboard. The checker automates this safely:

  • Reads the two cookies from your signed-in Chrome (in memory only).

  • Compares SHA-256 fingerprints of local vs. server-configured vs. server-persisted credentials.

  • Only on mismatch and an idle server, submits exactly one update through the real web form, then verifies persistence. Busy server → DEFERRED_BUSY, unresolved attempt → blocks further writes until confirmed.

  • Everything else (SSH tunnel, keys) is loopback-only; state files hold hashes, never credentials. Details: CREDENTIAL-CHECK.md.

npm test               # unit tests (no network)
npm run check          # one manual check
npm run check -- --read-only
cp launchagent/com.example.gemini2api-credential-check.plist \
   ~/Library/LaunchAgents/   # edit paths first
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.example.gemini2api-credential-check.plist

Coordinate calibration (for automation use)

Gemini's reported coordinates are not in the image's pixel space by default — the model guesses a coordinate system per answer (800×600, 1000×750, 711×711 observed on the same image). Declaring the exact dimensions in the prompt makes it pixel-accurate: 28/28 targets ≤1px across 4:3 / 2:1 / 16:9. Prompt recipe, conversion formula, and reproducible experiments: calibration/README.md.

Security notes

  • No credentials in code or logs: everything comes from .env (gitignored) or the environment; see .env.example.

  • The tunnel rejects unknown SSH hosts (RejectPolicy + known_hosts).

  • Cookie updates go through the dashboard's own form with masked inputs; the checker clears filled fields and the tunnel-origin management token afterwards.

  • service-patch/ additionally masks psid in admin API responses and runs the container as non-root.

License

MIT — except service-patch/, which derives from xwteam/gemini2api and follows its non-commercial license (personal study / research / self-deployment only). See service-patch/LICENSE-NOTE.md.

Available Tools

1 tool
gemini_chatGemini 单轮对话A

向 Gemini 发起一次独立对话。可附加本机图片、PDF、文本或 Office 文件的绝对路径;不保留上下文。默认使用 gemini-flash。

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo模型名,例如 gemini-flash 或 gemini-progemini-flash
imagesNo图片文件的本机路径
promptYes本轮要问 Gemini 的内容
systemNo可选的本轮系统指令
max_tokensNo可选的最大回答 token 数
attachmentsNoPDF、文本、表格、演示文稿等附件的本机路径

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses important behavior such as statelessness, attachment support, and default model, but omits operational details including authentication, rate limits, latency, and whether the call has side effects beyond sending content to Gemini.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and followed by key constraints. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter chat tool with full schema coverage and no output schema, the description covers the essential behavior: single-turn, stateless, optional local attachments, and default model. Remaining parameter details are delegated to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds only the absolute-path nuance for attachments and repeats file-type/default-model information already present in the schema; it does not explain system or max_tokens.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: initiate a one-off Gemini conversation (向 Gemini 发起一次独立对话) and clarifies it is stateless (不保留上下文), which distinguishes it from a multi-turn chat tool. Attachments and default model are also named, so an agent can identify the tool's scope immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No sibling tools exist, and the description supplies the key usage context: use it for an independent single-turn conversation, optionally with local attachments, and no prior context is kept. It does not include explicit when-not-to-use guidance, but the stateless single-turn scope is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedgemini_chat

TDQS

A4/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no tool-selection ambiguity. Its purpose is clearly stated as a single Gemini chat invocation.

Naming Consistency5/5

Only one tool, gemini_chat, uses a consistent snake_case style. There are no mixed naming conventions to evaluate.

Tool Count3/5

A single tool is thin for an API bridge and leaves little room for common Gemini operations. It is borderline rather than severely mismatched because the description scopes it to basic chat.

Completeness3/5

The tool covers single-turn chat with file attachments, but lacks multi-turn context, model selection, system instructions, and other common Gemini capabilities. These gaps limit the server to a narrow use case.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers