Claude-to-Gemini MCP Server
Claude-to-Gemini MCP 服务器
在 Claude Code 中将 Google Gemini 作为 MCP (Model Context Protocol) 服务器使用的 Agent-to-Agent 集成项目
🎯 项目目标
Claude Code: 主 AI(通用编码、调试、文件生成/修改)
Gemini: 辅助 AI(大规模上下文分析、代码库审查、6 种用途的图像生成)
Related MCP server: Claude Code Gemini MCP
✨ 主要功能
1. ask_gemini - 文本/代码生成
用途: 通用 Gemini 调用,大上下文分析
模型选择:
flash(默认): Gemini 2.5 Flash - 免费,快速pro: Gemini 3.1 Pro - 最新模型 (2026.02 发布),最高性能
上下文: 最大 1M token
2. gemini_analyze_codebase - 代码库分析
用途: 全面代码库专业分析
分析类型:
architecture: 架构模式分析duplications: 重复代码检测security: 安全漏洞检查performance: 性能优化机会general: 综合分析
3. generate_logo - 标志/图标生成
用途: 标志、图标、品牌资产制作
模型: Nano Banana Pro (
gemini-3-pro-image-preview) - 专注于专业资产制作参数:
prompt: 标志描述 (英文)brandName: 要包含的品牌/文本名称 (可选)style:minimal|modern|vintage|playful|corporatecolorScheme: 色系 (可选)
特点: 简洁且可缩放的设计,默认为 1:1 比例
4. generate_illustration - 插画/艺术作品生成
用途: 插画、艺术作品、角色、概念艺术
模型: Nano Banana 2 (
gemini-3.1-flash-image-preview) - 快速生成,免费参数:
prompt: 插画描述 (英文)style:watercolor|cartoon|vector|oil_painting|sketch|anime|pixel_artmood:cheerful|dark|calm|dramatic(可选)aspectRatio:1:1|16:9|9:16|4:3|3:4numberOfImages: 生成数量 (1-4)
5. generate_infographic - 信息图/图表生成
用途: 信息图、图表、流程图、时间轴
模型: Nano Banana Pro (
gemini-3-pro-image-preview) - 思考模式 + 文本渲染优化参数:
prompt: 信息图主题/内容 (英文)data: 要可视化的数据/信息 (可选)type:infographic|diagram|flowchart|timeline|comparison|statsaspectRatio:1:2(默认) |1:4|1:1|16:9
特点: 易读的文本渲染,支持纵向长布局
6. generate_photo - 写实照片生成
用途: 照片级真实图像、产品模型、广告摄影
模型: Imagen 4 (
imagen-4.0-generate-001) - 最高写实质量,付费参数:
prompt: 照片描述 (英文)style:natural|studio|cinematic|aerial|macronumberOfImages: 生成数量 (1-4)aspectRatio:1:1|16:9|9:16|4:3|3:4
特点: 最高 4K 分辨率,自动包含 SynthID 水印
7. generate_banner - 营销横幅/社交媒体图像生成
用途: 营销横幅、社交媒体图像、缩略图、海报
模型: Nano Banana Pro (
gemini-3-pro-image-preview) - 文本 + 图形组合参数:
prompt: 横幅描述 (英文)text: 要包含在横幅中的文本 (可选)platform:facebook|instagram|twitter|youtube|linkedin|webaspectRatio: 按平台自动设置
特点: 提供各平台的最佳尺寸预设
8. edit_image - 图像编辑/修改
用途: 在现有图像中添加/删除/修改元素
模型: Nano Banana 2 (
gemini-3.1-flash-image-preview) - 免费参数:
prompt: 编辑说明 (英文)imagePath: 要编辑的图像文件路径action:modify|add|remove|style_transfer|enhance
特点: 支持交错编辑、多轮对话式修改
🛠 技术栈
Runtime: Node.js 18+
MCP SDK: @modelcontextprotocol/sdk
AI API: Google Gemini API (@google/generative-ai)
IDE: Claude Code (CLI + VSCode 扩展)
📦 安装方法
1. 前置准备
安装 Node.js 18 或更高版本
订阅 Claude Pro/Max 计划
获取 Google Gemini API 密钥 (ai.google.dev)
2. 克隆项目
git clone https://github.com/YOUR_USERNAME/claude-to-gemini.git
cd claude-to-gemini3. 安装依赖
npm install4. 注册 MCP 服务器
claude mcp add gemini \
--env GEMINI_API_KEY=YOUR_API_KEY_HERE \
-- node /ABSOLUTE_PATH/claude-to-gemini/index.js注意:
将
YOUR_API_KEY_HERE替换为实际的 Gemini API 密钥将
/ABSOLUTE_PATH/替换为实际的项目路径 (例如:/Users/username/projects/claude-to-gemini/index.js)
5. 验证
claude mcp list输出示例:
gemini - node /Users/username/projects/claude-to-gemini/index.js🚀 使用方法
启动 Claude Code
claude基本使用 (Flash 模型,免费)
ask_gemini 도구를 사용해서 "이 프로젝트 전체 구조를 분석해줘" 물어봐줘使用 Pro 模型 (付费,高性能)
ask_gemini 도구를 사용해서 model을 "pro"로 설정하고 "복잡한 아키텍처 설계해줘" 물어봐줘代码库分析
gemini_analyze_codebase 도구로 보안 취약점을 찾아줘标志生成
generate_logo 도구로 brandName을 "CafeKiosk"로, style을 "modern"으로 설정하고
"A minimalist coffee cup logo with geometric shapes" 로고 만들어줘插画生成 (免费)
generate_illustration 도구로 style을 "watercolor"로 설정하고
"A cozy cafe interior with warm lighting" 삽화 생성해줘信息图生成
generate_infographic 도구로 type을 "flowchart"로 설정하고
"User authentication flow: login, verify, 2FA, dashboard" 다이어그램 만들어줘写实照片生成 (付费 - Imagen 4)
generate_photo 도구로 style을 "studio"로, numberOfImages를 4로 설정하고
"Professional food photography of a latte with beautiful latte art" 이미지 4개 생성해줘营销横幅生成
generate_banner 도구로 platform을 "instagram"으로 설정하고
text를 "Grand Opening 50% OFF"로
"Bright modern cafe promotion banner with coffee beans" 배너 만들어줘图像编辑 (免费)
edit_image 도구로 action을 "remove"로, imagePath를 "./photo.png"으로 설정하고
"Remove the background person and keep only the coffee cup" 편집해줘💡 使用场景
场景 1: 新项目架构设计
ask_gemini 도구로 React + Express + PostgreSQL
전자상거래 앱의 전체 아키텍처를 설계해줘场景 2: 遗留代码分析
gemini_analyze_codebase 도구로
focus를 'duplications'로 설정해서 중복 코드를 찾아줘场景 3: 大规模重构
ask_gemini 도구로 이 프로젝트 전체를 읽고
모던한 아키텍처로 마이그레이션 계획을 세워줘📚 实战指南
如何在工作中应用?
更多详细的实战应用方法,请参考 📖 实战应用指南 (USECASES.md)!
主要内容:
🔍 助手代码审查 (每日早晨例行)
🏗️ 大规模重构 (1200 行迁移)
🚀 项目入职 (1 小时内掌握核心)
🎨 架构设计 (Monorepo 结构)
🖼️ 各类图像生成 (标志、插画、信息图、照片、横幅、编辑)
💡 提示与技巧 (成本优化、模型选择)
📊 模型对比
文本/代码生成模型
模型 | 上下文 | 成本 | 速度 | 推荐用途 |
Gemini 2.5 Flash | 1M token | 免费 | 快速 | 通用分析,大多数任务 |
Gemini 3.1 Pro | 1M token | 付费 | 快速 | 最高性能,复杂推理 |
图像生成模型
模型 | 工具 | 用途 | 成本 | 特点 |
Nano Banana Pro ( |
| 标志、信息图、横幅 | 付费 | 专业资产制作、文本渲染、思考模式 |
Nano Banana 2 ( |
| 插画、图像编辑 | 免费 | 快速生成、对话式编辑、多种画风 |
Imagen 4 ( |
| 写实照片、产品模型 | 付费 | 最高 4K、照片级真实、包含 SynthID |
⚠️ 安全注意事项
API 密钥保护
严禁:
❌ 将 API 密钥上传到 GitHub
❌ 在代码中硬编码 API 密钥
❌ 在公共场所分享 API 密钥
建议:
✅ 仅通过环境变量管理
✅ 在
.gitignore中包含.claude.json✅ API 密钥泄露时立即重新生成
.gitignore 必备内容
node_modules/
.claude.json
.env
*.key🤝 贡献方式
Fork 项目
创建功能分支 (
git checkout -b feature/AmazingFeature)提交更改 (
git commit -m 'Add some AmazingFeature')推送到分支 (
git push origin feature/AmazingFeature)开启 Pull Request
📝 许可证
MIT License - 详情请参阅 LICENSE 文件
🔗 参考资料
📧 联系方式
项目相关咨询: GitHub Issues
Made with ❤️ by [Your Name]
Available Tools
4 toolsask_geminiA
Use Gemini for large context analysis (1M tokens), architecture design, or whole codebase review. Best for tasks requiring understanding of entire projects.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | The question or task for Gemini | |
| context | No | Optional: Large codebase, multiple files, or extensive context to analyze | |
| model | No | Model to use: 'flash' (default, free, fast) or 'pro' (3 Pro, latest model, better quality, paid) | flash |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the 1M token capacity and model options (free/fast vs paid/better quality), which adds useful context about capabilities and cost implications. However, it doesn't cover rate limits, error handling, response format, or authentication requirements that would be important for a tool like this.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with two sentences that each earn their place. The first sentence establishes the core purpose and key differentiators, while the second provides the essential usage guidance. No wasted words or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (AI model interaction with large context), lack of annotations, and no output schema, the description is adequate but has clear gaps. It covers the main use cases and capacity but doesn't address response format, error conditions, or operational constraints that would be important for complete understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema. The baseline of 3 is appropriate when the schema does the heavy lifting for parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('use Gemini for large context analysis, architecture design, or whole codebase review') and distinguishes it from siblings by emphasizing its suitability for tasks requiring understanding of entire projects, unlike image generation tools or potentially more focused code analysis tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('large context analysis, architecture design, or whole codebase review'), but doesn't explicitly state when NOT to use it or name specific alternatives among the sibling tools. It implies usage for extensive tasks but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_analyze_codebaseC
Specialized tool for analyzing entire codebases. Gemini will find patterns, duplications, architectural issues, and suggest improvements.
| Name | Required | Description | Default |
|---|---|---|---|
| codebase | Yes | The entire codebase or multiple files concatenated | |
| focus | No | What to focus on: 'architecture', 'duplications', 'security', 'performance', or 'general' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the tool 'will find patterns, duplications, architectural issues, and suggest improvements,' but lacks details on how it operates (e.g., processing time, output format, limitations like codebase size, or whether it modifies code). For a complex analysis tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two sentences that efficiently state the tool's purpose and capabilities. It's front-loaded with the main function ('analyzing entire codebases') and avoids unnecessary details. However, it could be slightly more structured by explicitly separating scope from outcomes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of codebase analysis, lack of annotations, and no output schema, the description is incomplete. It doesn't cover behavioral aspects like processing constraints, error handling, or result format, which are crucial for an AI agent to use the tool effectively. The description should compensate for these gaps but falls short.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters ('codebase' and 'focus') with descriptions and an enum for 'focus'. The description adds no additional meaning beyond what the schema provides, such as explaining how the 'codebase' should be formatted or what 'general' focus entails. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'analyzing entire codebases' with specific outcomes like finding patterns, duplications, architectural issues, and suggesting improvements. It uses specific verbs ('find', 'suggest') and identifies the resource ('codebases'), but doesn't explicitly differentiate from sibling tools like 'ask_gemini' which might also handle code analysis in a different way.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'ask_gemini' (which might handle general queries) or specify contexts where this specialized analysis is preferred over other options. Usage is implied by the description but lacks explicit when/when-not instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_image_geminiB
Generate images using Gemini 2.5 Flash Image (Nano Banana). Best for contextual understanding, image editing, multi-image composition, and iterative refinement. Free tier available.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Description of the image to generate (in English, max 480 tokens) | |
| numberOfImages | No | Number of images to generate (1-4, default: 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the model name ('Gemini 2.5 Flash Image (Nano Banana)') and use cases, but doesn't disclose important behavioral traits like rate limits, authentication needs, cost implications beyond 'Free tier available', or what happens on failure. The free tier mention is useful but insufficient for full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise with two sentences that each serve a purpose: the first states the core function and model, the second provides usage context and cost information. It's front-loaded with the main purpose. However, the parenthetical model name '(Nano Banana)' adds minor clutter without clear value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 2 parameters with 100% schema coverage but no annotations and no output schema, the description is moderately complete. It covers the what and some when, but lacks important context about behavioral constraints, error handling, and output format. For an image generation tool with potential cost/rate implications, more completeness would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description doesn't add any meaningful parameter semantics beyond what's in the schema - it doesn't explain prompt best practices, token limitations beyond the schema's 'max 480 tokens', or how 'numberOfImages' affects output. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates images using a specific AI model (Gemini 2.5 Flash Image), which is a specific verb+resource combination. It distinguishes from sibling tools like 'ask_gemini' and 'gemini_analyze_codebase' by focusing on image generation rather than text analysis or code review. However, it doesn't explicitly differentiate from 'generate_image_imagen', which appears to be a similar image generation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some context about when to use this tool ('Best for contextual understanding, image editing, multi-image composition, and iterative refinement'), which implies usage scenarios. However, it doesn't explicitly state when NOT to use it or mention alternatives like the sibling 'generate_image_imagen' tool, leaving the agent to infer the best choice between similar tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_image_imagenB
Generate images using Imagen 4. Best for photorealistic quality, high-resolution outputs, and professional branding. Paid service.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Description of the image to generate (in English, max 480 tokens) | |
| numberOfImages | No | Number of images to generate (1-4, default: 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions 'Paid service' (implying cost/access restrictions) and quality aspects, but lacks critical behavioral details: it doesn't specify rate limits, authentication needs, output format (e.g., image URLs or files), processing time, or error handling. For a generative AI tool with no annotation coverage, this leaves significant gaps in understanding operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and well-structured in a single sentence, with no wasted words. It front-loads the core action ('Generate images using Imagen 4') and efficiently lists key features and constraints, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of an image generation tool with no annotations and no output schema, the description is incomplete. It lacks information on output format (e.g., how images are returned), error conditions, cost details beyond 'Paid service,' and comparison with sibling tools. For a tool that likely produces binary or URL outputs, this omission is significant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents both parameters (prompt and numberOfImages). The description adds no parameter-specific information beyond what's in the schema, such as prompt best practices or image count implications. Baseline 3 is appropriate as the schema handles parameter documentation adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Generate images using Imagen 4' with specific capabilities ('photorealistic quality, high-resolution outputs, professional branding'). It distinguishes from sibling tools by specifying the Imagen 4 model, but doesn't explicitly contrast with 'generate_image_gemini' beyond mentioning 'Paid service' versus likely free alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage context with 'Best for photorealistic quality...' and 'Paid service,' which implies when to prefer this over free alternatives. However, it doesn't explicitly state when to use this versus 'generate_image_gemini' or other siblings, nor does it mention any prerequisites or exclusions beyond the cost implication.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- Changed
ask_gemini1 field changed- changed
Input schema / properties / model / descriptionPrevious value: -"Model to use: 'flash' (default, free, fast) or 'pro' (2.5 Pro, 1M tokens, better quality, paid)"New value: +"Model to use: 'flash' (default, free, fast) or 'pro' (3 Pro, latest model, better quality, paid)"
- Added
generate_image_gemini - Added
generate_image_imagen
2 tool updates
- First observed
ask_gemini - First observed
gemini_analyze_codebase
TDQS
Scored across 4 tools
The tools have overlapping purposes that could cause confusion. 'ask_gemini' and 'gemini_analyze_codebase' both target Gemini for analysis tasks, with the latter being a specialized subset of the former. The two image generation tools are clearly distinct in their use cases (Gemini for contextual/iterative work, Imagen for photorealism), but the analysis tools are not well-differentiated.
Naming conventions are inconsistent. 'ask_gemini' uses a verb-object pattern, 'gemini_analyze_codebase' uses a noun-verb-object pattern with underscores, and both image tools use 'generate_image_' prefix but with different suffixes ('gemini' vs 'imagen'). This mixed style lacks a predictable pattern.
Four tools is a reasonable count for a server bridging Claude and Gemini/Imagen services. It covers analysis and image generation without being overly sparse or bloated. However, the scope feels slightly thin given the potential breadth of interactions between these AI systems.
The server covers text analysis and image generation but has notable gaps. There are no tools for conversational interactions, file processing, or multimodal tasks beyond image generation. The domain appears to be 'Claude-to-Gemini integration,' but the surface lacks tools for common workflows like chat, document analysis, or combined text-image tasks.
Maintenance
Related MCP Connectors
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Claude Code / MCP skills for the dev pipeline: discover, spec, design, build, ship, operate.
MCP server unifying ERPs, CRMs, APIs and knowledge base for Claude, ChatGPT and Gemini.
Related MCP Servers
- FlicenseBqualityNot gradedmaintenanceEnables Claude Code to use Google Gemini AI capabilities for analyzing PDFs and images, generating and translating text, and reviewing code. Supports both CLI and API backends with different quota limits.481 npm-
- FlicenseCqualityDmaintenanceEnables Claude Code to call Gemini models through an OpenAI-compatible API, providing tools for deep analysis, brainstorming, code review, and general queries using Gemini's capabilities.4-
- AlicenseBqualityDmaintenanceEnables Claude to collaborate with Gemini for code reviews, second opinions, and iterative software development. It facilitates multi-step workflows including PRD creation and code generation through an AI orchestration framework.28 npm1MIT
- AlicenseAqualityCmaintenanceIntegrates Google's Gemini AI models into Claude Code and other MCP clients to provide second opinions, code comparisons, and token counting. It supports streaming responses and multi-turn conversations directly within your existing AI development workflow.3Apache 2.0