gemini2api-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gemini2api-mcpsummarize this PDF for me: /Users/me/docs/report.pdf"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
gemini2api-mcp
MCP wrapper, credential auto-sync, and service patches for gemini2api — a self-hosted service that turns the Gemini web UI into an OpenAI-compatible API using browser cookies.
This repository contains everything that runs on the client side of such a deployment:
Component | What it does |
| MCP server exposing a |
| Credential auto-sync: keeps the server's Google cookies fresh by reading the signed-in Chrome session, comparing SHA-256 fingerprints via the admin API, and submitting at most one idle-window web-form update |
| macOS LaunchAgent template that runs the checker every 5 minutes |
| Hardening and bug-fix patches for the upstream gemini2api service (non-root Docker, masked credentials in admin API, idle-reservation cookie updates, atomic persistence + regression test) |
| Research: how Gemini reports pixel coordinates, and the prompt recipe that makes them pixel-accurate (28/28 targets ≤1px) |
Required dependencies / upstream projects
xwteam/gemini2api (upstream service, non-commercial license) The server this repo wraps. Deploy it first; it provides the
http://host:5918/openai/v1endpoint, the/adminmanagement API, and the cookie-based account pool. The files underservice-patch/are derivatives of this project and follow its license.MCP TypeScript SDK (
@modelcontextprotocol/server, npm) — MCP protocol implementation used byserver.mjs.zod (npm) — tool input schemas.
paramiko (Python) — loopback SSH tunnel for management traffic (
ssh_tunnel.py).A Chrome automation bridge speaking the
{action, args}JSON protocol on127.0.0.1:10086(navigate / fill / click / evaluate / cdpNetwork.getCookies). The reference setup uses the Kimi WebBridge extension + native daemon (~/.kimi-webbridge/bin/kimi-webbridge); any CDP-based bridge with the same command shape works.Node.js ≥ 20 (built-in
fetch,AbortSignal.timeout) and Python ≥ 3.11. The scheduler/notifier part is macOS-only (LaunchAgent +osascript); the MCP wrapper itself is cross-platform.
Related: UI-Venus-MCP — a
cross-platform computer-use MCP that can use this endpoint as a pluggable
vision provider (see its docs/gemini-coordinate-calibration.md).
Related MCP server: OpenAI-Compatible MCP Gateway
Quick start
git clone https://github.com/q1820926174-cpu/gemini2api-mcp
cd gemini2api-mcp
npm install
cp .env.example .env # fill in your server URL + keys
npm start # stdio MCP serverRegister in an MCP client (ZCode ~/.zcode/cli/config.json example):
{
"mcp": { "servers": {
"gemini2api": {
"command": "node",
"args": ["/absolute/path/to/gemini2api-mcp/server.mjs"]
}
}}
}Tool surface (gemini_chat):
{
"prompt": "What is in this image?",
"model": "gemini-flash", // or gemini-pro / -lite / -thinking
"system": "optional system instruction",
"images": ["/abs/path/pic.png"], // up to 8, image types
"attachments": ["/abs/path/doc.pdf"], // up to 8, pdf/text/office
"max_tokens": 4096
}Limits: single-round only (no conversation state), max 8 files / 20 MiB per call, 180 s upstream timeout (image calls take 5–120 s; MCP client timeouts below ~180 s will cut image requests short).
Credential auto-sync (optional)
Google rotates the __Secure-1PSID* cookies behind gemini2api; when that
happens the service starts failing until someone re-pastes cookies into the
admin dashboard. The checker automates this safely:
Reads the two cookies from your signed-in Chrome (in memory only).
Compares SHA-256 fingerprints of local vs. server-configured vs. server-persisted credentials.
Only on mismatch and an idle server, submits exactly one update through the real web form, then verifies persistence. Busy server →
DEFERRED_BUSY, unresolved attempt → blocks further writes until confirmed.Everything else (SSH tunnel, keys) is loopback-only; state files hold hashes, never credentials. Details: CREDENTIAL-CHECK.md.
npm test # unit tests (no network)
npm run check # one manual check
npm run check -- --read-only
cp launchagent/com.example.gemini2api-credential-check.plist \
~/Library/LaunchAgents/ # edit paths first
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.example.gemini2api-credential-check.plistCoordinate calibration (for automation use)
Gemini's reported coordinates are not in the image's pixel space by default — the model guesses a coordinate system per answer (800×600, 1000×750, 711×711 observed on the same image). Declaring the exact dimensions in the prompt makes it pixel-accurate: 28/28 targets ≤1px across 4:3 / 2:1 / 16:9. Prompt recipe, conversion formula, and reproducible experiments: calibration/README.md.
Security notes
No credentials in code or logs: everything comes from
.env(gitignored) or the environment; see.env.example.The tunnel rejects unknown SSH hosts (
RejectPolicy+known_hosts).Cookie updates go through the dashboard's own form with masked inputs; the checker clears filled fields and the tunnel-origin management token afterwards.
service-patch/additionally maskspsidin admin API responses and runs the container as non-root.
License
MIT — except service-patch/, which derives from
xwteam/gemini2api and follows its
non-commercial license (personal study / research / self-deployment only).
See service-patch/LICENSE-NOTE.md.
Available Tools
1 toolgemini_chatGemini 单轮对话A
向 Gemini 发起一次独立对话。可附加本机图片、PDF、文本或 Office 文件的绝对路径;不保留上下文。默认使用 gemini-flash。
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | 模型名,例如 gemini-flash 或 gemini-pro | gemini-flash |
| images | No | 图片文件的本机路径 | |
| prompt | Yes | 本轮要问 Gemini 的内容 | |
| system | No | 可选的本轮系统指令 | |
| max_tokens | No | 可选的最大回答 token 数 | |
| attachments | No | PDF、文本、表格、演示文稿等附件的本机路径 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses important behavior such as statelessness, attachment support, and default model, but omits operational details including authentication, rate limits, latency, and whether the call has side effects beyond sending content to Gemini.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and followed by key constraints. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter chat tool with full schema coverage and no output schema, the description covers the essential behavior: single-turn, stateless, optional local attachments, and default model. Remaining parameter details are delegated to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds only the absolute-path nuance for attachments and repeats file-type/default-model information already present in the schema; it does not explain system or max_tokens.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: initiate a one-off Gemini conversation (向 Gemini 发起一次独立对话) and clarifies it is stateless (不保留上下文), which distinguishes it from a multi-turn chat tool. Attachments and default model are also named, so an agent can identify the tool's scope immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No sibling tools exist, and the description supplies the key usage context: use it for an independent single-turn conversation, optionally with local attachments, and no prior context is kept. It does not include explicit when-not-to-use guidance, but the stateless single-turn scope is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
gemini_chat
TDQS
Scored across 1 tool
Only one tool exists, so there is no tool-selection ambiguity. Its purpose is clearly stated as a single Gemini chat invocation.
Only one tool, gemini_chat, uses a consistent snake_case style. There are no mixed naming conventions to evaluate.
A single tool is thin for an API bridge and leaves little room for common Gemini operations. It is borderline rather than severely mismatched because the description scopes it to basic chat.
The tool covers single-turn chat with file attachments, but lacks multi-turn context, model selection, system instructions, and other common Gemini capabilities. These gaps limit the server to a narrow use case.
Maintenance
Related MCP Connectors
Remote streamable-HTTP MCP server running on a single Cloudflare Worker. Your assistant gets live Airbnb, Amazon, Booking.com, Google Flights, Maps and Reddit data, social search on X, Instagram and TikTok, the Meta Ad Library, and image/video generation without any keys. Connect your own accounts to let it send WhatsApp or Telegram messages, work an IMAP inbox, manage Meta Ads campaigns and publish to X and LinkedIn. OAuth 2.1 with PKCE; stored credentials are AES-256-GCM encrypted.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Zero-setup MCP gateway securely connecting AI to your tools with authentication and workflows
Give any MCP-compatible AI assistant a builder for live, hosted web tools and workflows.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceIntegrates Google's Gemini API with MCP-compatible clients for chat, real-time web search, knowledge queries, code/text analysis, and content generation.107 npmMIT
- FlicenseNot gradedqualityDmaintenanceLocal MCP server that exposes fixed tools for GPT, Claude, and Gemini while routing to any OpenAI-compatible chat completions backend with independent configuration per target.1-
- AlicenseAqualityDmaintenanceMCP server for Google's Gemini API, enabling text, image, video, speech, embeddings, and deep research capabilities through a single tool set.10MIT
- AlicenseNot gradedqualityDmaintenanceRemote MCP server that exposes Google Gemini's text, image, video (Veo), and audio transcription capabilities as tools any MCP client can call directly.91 npmMIT