Skip to main content
Glama

CLIProxy MCP

把 CLIProxy 反代出的 ChatGPT 生图能力,接入 Codex、Claude Code 和 ZCode。

主要面向已经通过 CLIProxy 暴露 ChatGPT 生图接口的用户:在 AI 编程客户端中直接生成文章封面、正文配图、社交媒体配图或项目插画,并把原图保存到本地。

Codex / Claude Code / ZCode
          ↓ MCP
     CLIProxy MCP
          ↓ /v1/images/generations
  CLIProxy → ChatGPT 生图能力
          ↓
     本地原图 + 客户端预览

CLIProxy 负责上游账号接入和反代,本项目负责 MCP 工具封装、客户端配置以及图片保存。只需要提供 CLIProxy 的 API 地址和访问密钥,无需在此项目中另填 ChatGPT 登录凭据或配置文本模型。

支持 gpt-image-2.5-flare(默认)、gpt-image-2.5-sunburst,也可填写你的 CLIProxy 实际提供的其他兼容图片模型 ID。这些名称按代理暴露的模型 ID 使用,可用性以你的服务为准。

版本说明:快速向导和本地直启从 0.2.1 起提供。旧版用户请重新运行新版 init,更新客户端启动配置。

快速开始

1. 准备环境

  • 安装 Node.js 22.13 或更新版本,确认终端可以运行 node 和 npx。

  • 在 CLIProxy 中完成 ChatGPT 上游接入,确认已提供可用的生图模型,以及兼容的 POST /v1/images/generations 接口。

  • 准备 CLIProxy 的 API 地址和访问密钥,例如 http://127.0.0.1:8317/v1。

  • CLIProxy 必须能从运行 MCP 的电脑访问。远程或云端环境的 127.0.0.1 不指向你的电脑。

2. 运行配置向导

npx -y cliproxy-mcp@latest init

向导会依次完成:

  1. 输入 API 地址和密钥,密钥默认隐藏;已有密钥可留空保留。

  2. 查询模型列表验证连接,鉴权失败可直接修改后重试。

  3. 选择“快速配图”:首次配置优先使用服务提供的 flare / sunburst,沿用默认图片目录;再次运行保留已有设置。只需要生图时,无需设置文本模型。 要切换图片模型、修改目录或通用导出时,选择“自定义设置”。

  4. 多选 Codex、Claude Code、ZCode;已有接入会预选,检测到配置的客户端会标注。

  5. 查看易读摘要,一次确认全部变更。向导准备运行环境、验证 MCP 握手和工具发现,再完成客户端接入。

连接验证只查询模型列表,不会自动生成文本或图片。

3. 在客户端中使用

重新加载客户端的 MCP 服务;如果没有出现 cliproxy,重启客户端。然后试着说:

用 CLIProxy 为这篇文章生成一张横版封面,简洁插画风格,不要文字。

用 gpt-image-2.5-sunburst 生成一张产品介绍配图,并告诉我原图保存在哪里。

也可以先说“调用 cliproxy 查看可用模型”,确认代理暴露的模型 ID。

图片保存到本地,工具返回 原图绝对路径、实际尺寸和预览。默认输出目录为 ~/.cliproxy-mcp/output;可以将生成的图片继续用于文章、页面或项目素材。

Related MCP server: smartapi-images

命令

命令

功能

npx -y cliproxy-mcp@latest init

首次配置、修改设置、接入更多客户端

npx -y cliproxy-mcp@latest doctor

检查 npx、配置、API 连接、目录权限和客户端配置

npx -y cliproxy-mcp@latest config show

查看环境变量覆盖后的生效配置,隐藏密钥

npx -y cliproxy-mcp@latest --version

查看版本

npx -y cliproxy-mcp@latest --help

查看帮助

使用独立配置文件:

npx -y cliproxy-mcp@latest init --config /absolute/path/cliproxy.json
npx -y cliproxy-mcp@latest doctor --config /absolute/path/cliproxy.json

init 需要交互终端。Ctrl+C 可退出,已经确认保存的步骤会保留。MCP 客户端启动参数中不要填写 init:无子命令时才启动 stdio MCP 服务。服务由客户端负责启动,无需另外运行常驻终端,也没有 HTTP 监听端口。

doctor 检查引用的配置文件及本地运行入口是否存在,不发起生成请求,失败时返回非零退出码。它检查客户端配置文件,不替代客户端内的实际连接检查。

全局安装与升级

也可使用全局安装:

npm install -g cliproxy-mcp@latest
cliproxy-mcp init

升级后重新运行新版向导,并选择已接入的客户端:

npx -y cliproxy-mcp@latest init

向导会把当前版本安装到 ~/.cliproxy-mcp/runtime/<版本>/,然后让客户端通过 Node.js 的绝对路径直接启动,握手期间不再运行 npm 或访问 registry。运行环境安装与校验成功后才更新客户端配置;重新运行可复用已安装版本或修复损坏安装。

客户端配置固定本次运行的包版本,不会在每次启动时自动升级。重新运行新版向导会准备新版本并更新启动路径,接口配置和图片不会因包升级被覆盖。可用 cliproxy-mcp@0.2.2 固定该版本。

配置与客户端接入

一份接口配置

默认配置保存在 ~/.cliproxy-mcp/config.json(~ 表示用户主目录)。客户端只通过 CLIPROXY_CONFIG 引用文件路径,不重复复制密钥。

{
  "baseUrl": "http://127.0.0.1:8317/v1",
  "apiKey": "YOUR_API_KEY",
  "imageModel": "gpt-image-2.5-flare",
  "outputDir": "./output",
  "timeoutMs": 300000,
  "allowedImageOrigins": []
}
  • baseUrl 包含 API 前缀(一般是 /v1),不包含 /images/generations 等具体接口路径。

  • imageModel 是 CLIProxy 暴露的图片模型 ID。上面的配置即可用于生图,不需要文本模型。

  • 相对 outputDir 按配置文件所在目录解析;默认用户配置对应 ~/.cliproxy-mcp/output。

  • 可通过 --config 或 CLIPROXY_CONFIG 指定其他文件。

  • 默认先读用户配置;不存在时,源码开发模式兼容读取项目根目录的 config.json。安装包不包含真实配置。

  • 密钥保存在本机配置文件中,不是加密存储;分享项目时使用 config.example.json,不要分享真实配置或备份。

也支持完全通过环境变量配置。以下环境变量会覆盖文件值:

环境变量

对应字段

CLIPROXY_BASE_URL

baseUrl

CLIPROXY_API_KEY

apiKey

CLIPROXY_IMAGE_MODEL

imageModel

CLIPROXY_OUTPUT_DIR

outputDir

CLIPROXY_TIMEOUT_MS

timeoutMs

向导验证并保存输入值;如果保留了上述环境变量,运行服务时它们仍会覆盖文件。程序不自动加载 .env。

客户端配置位置

目标

用户级文件

MCP 配置键

Codex

~/.codex/config.toml,支持 CODEX_HOME

mcp_servers.cliproxy

Claude Code

~/.claude.json;设置 CLAUDE_CONFIG_DIR 时为该目录下的 .claude.json

mcpServers.cliproxy

ZCode

~/.zcode/cli/config.json

mcp.servers.cliproxy

通用导出

~/.agents/mcp.json

mcpServers.cliproxy

向导保留其他设置和 MCP 服务,修改前在同目录生成 .bak 备份。重复运行且设置一致时不写入;损坏配置、锁冲突或预览后文件变化会阻止对应写入。Codex 仅修改 cliproxy 对应段,保留其他段的原文和注释;特殊内联 TOML 写法无法安全修改时会停止,避免整份重写。JSON 文件需为标准 JSON,暂不支持 JSONC。

.agents/mcp.json 仅在“自定义设置”中提供,是可选通用导出,不是三个客户端的统一读取入口。 Codex 和 Claude Code 使用原生文件;ZCode 在同一作用域有原生 MCP 服务时会忽略 .agents。首次为 ZCode 创建原生服务时,向导会把已有 .agents 服务一起保留到原生配置,并提示数量,不修改原 .agents 文件。

客户端配置使用本机 Node.js 可执行文件和运行入口的绝对路径,不经过 shell;接口配置路径放在环境变量中。Node.js 安装位置变化或运行目录被删除后,重新运行 init 即可修复。Codex 启动超时设置为 120 秒、工具超时为 960 秒;其他客户端的工具超时需按其设置调整。

快速流程示意

API 地址 → API 密钥 → 快速配图 → 选择客户端 → 确认应用

即将应用
  图片模型:gpt-image-2.5-flare
  图片目录:用户主目录/.cliproxy-mcp/output
  Codex:添加 → 用户主目录/.codex/config.toml
  ZCode:添加 → 用户主目录/.zcode/cli/config.json

✓ 接口配置已保存
✓ MCP 启动正常:3 个工具,显示本次实际耗时
✓ Codex:已接入
✓ ZCode:已接入
下一步:在所选客户端重新加载 MCP,或重启客户端。

只有在有需要时才显示文本模型、手动模型、目录修改和通用导出选项。安装或握手失败时不会宣称客户端接入成功;已经保存的接口配置会保留,方便下次重试。

配置规范参考:Codex MCP、Claude Code MCP、ZCode MCP。

生图工具与模型选择

工具

用途

CLIProxy 接口

generate_image

生成一张图片,保存原图并返回尺寸和预览

POST /images/generations

list_models

查询代理暴露的模型 ID,辅助排查配置

GET /models

模型 ID

如何选择

gpt-image-2.5-flare

默认选项;快速配图优先使用代理列表中的此模型

gpt-image-2.5-sunburst

可在提示词、工具参数或配置中指定

其他 ID

使用 CLIProxy 实际暴露且支持生图接口的模型 ID

本项目不对这些模型的画质、速度或额度作固定比较,具体表现取决于代理及上游。模型列表出现某个 ID 也不等于它一定支持生图;init 和 doctor 的连接检查不会替你发起生成请求。

generate_image 参数示例:

{
  "prompt": "一只橙色机器人在绿植旁读书,简洁插画,奶油色背景,无文字",
  "model": "gpt-image-2.5-flare",
  "size": "1024x1024"
}

model 可省略,使用配置中的默认图片模型;size 默认 1024x1024,也可请求其他尺寸。上游可能返回不同尺寸,以工具结果中的实际宽高为准。每次调用生成一张图片,需要多张时分别调用。

可选:文本生成

generate_text 是附带的兼容工具,不影响生图,首次配图时无需设置默认文本模型。只有需要通过同一 CLIProxy 另外调用文本模型时,再选择“配图 + 文本生成”。

它支持 Chat Completions 和 Responses。可在调用时传 model,或在配置中设置 textModel、textApi;对应环境变量为 CLIPROXY_TEXT_MODEL、CLIPROXY_TEXT_API。

{
  "model": "YOUR_TEXT_MODEL_ID",
  "prompt": "用一句话解释 MCP",
  "system": "请用中文简洁回答",
  "api": "chat",
  "maxTokens": 256
}

api 可改为 responses。maxTokens 分别映射到 max_completion_tokens 和 max_output_tokens;代理不支持时可省略。

常见问题

现象

处理方式

HTTP 401

检查填写的是 CLIProxy 的访问密钥,并确认没有旧环境变量覆盖配置

连接检查成功,但生图失败

检查 CLIProxy 的 ChatGPT 上游状态、模型 ID 和 images/generations 支持情况;模型列表查询成功不代表生成接口一定可用

HTTP 404

检查地址中的 /v1、接口类型和模型 ID

HTTP 429

检查上游额度或速率限制,稍后手动重试

找不到 npx

安装 Node.js/npm;确认客户端启动环境的 PATH,安装后重启客户端

首次连接超时,稍后重试才正常

升级后重新运行 init 并选择客户端,将旧的 npx 启动配置替换为本地直接启动

客户端没有出现工具

确认向导中选中了该客户端并写入成功,重新加载 MCP 或重启客户端

直接运行无输出

无子命令是 stdio 服务入口,会等待客户端输入;首次配置请运行 init

init 提示需要交互终端

在 PowerShell、Terminal 等终端运行,不要放进 MCP 服务启动参数

图片下载来源未允许

将可信图片 CDN 来源加入 allowedImageOrigins,例如 https://images.example.com

超时或取消

上游可能仍在执行或已经计费;确认后再重试,避免重复生成

先运行 doctor 排查配置和连接,再查看客户端的 MCP 状态。

行为与限制

  • 生成请求可能消耗上游额度,失败不自动重试。服务端默认超时 300 秒,可配置 1–900 秒;客户端超时也需足够长。

  • 每次生成一张图片,接受 Base64 或 URL 响应。保存原图,同时提供最长边 768 像素的 PNG 预览。

  • 支持 PNG/JPEG/WebP,单图最多 20 MiB、4000 万像素,JSON 响应最多 32 MiB。请求尺寸只传给上游,实际尺寸以结果为准。

  • 图片 URL 默认仅允许 API 同源,可用 allowedImageOrigins 添加可信来源。下载不发送 API 密钥、不跟随重定向;需要鉴权的下载 URL 暂不支持。

  • 错误不会透传上游响应正文。远程客户端可能无法打开本机原图路径,但可查看 MCP 预览。

  • 当前支持单轮文本和文生图,未包含图生图、多轮会话托管、流式返回、异步任务恢复或 MCP HTTP 传输。

开发与测试

在源码根目录执行:

npm ci
npm run check
npm test
node src/index.js init
node src/index.js doctor

测试覆盖真实 stdio MCP 握手、工具调用、图片校验、401 脱敏、配置向导、客户端合并、备份及并发修改保护。使用模拟 HTTP 服务,不需要真实密钥或付费调用。

已有配置后,可通过真实 MCP 客户端进行只读连通性检查:

node scripts/smoke.js

源码根目录的同名 npm 项目可能影响 npx 命令解析,开发时使用 node src/index.js。

打包:

npm pack

prepack 自动执行语法检查和测试。通过文件白名单,安装包只包含运行代码、配置示例、README、MIT 许可证和包元数据,不包含密钥、真实配置、备份或生成图片。.tgz 可通过 npm install -g /path/to/cliproxy-mcp-0.2.2.tgz 安装。

许可证

MIT © 2026 bruc3van。

Available Tools

3 tools
generate_imageA

通过 CLIProxy 的 images/generations 生成一张图片并保存本地,返回绝对路径、实际尺寸和 PNG 预览。默认 gpt-image-2.5-flare,也可指定 gpt-image-2.5-sunburst 或其他模型。可能消耗额度,失败不自动重试。

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNo请求尺寸,例如 1024x1024、1536x1024;支持情况取决于模型,实际尺寸以结果为准
modelNoCLIProxy 模型 ID;省略时使用配置中的默认模型
promptYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only state not read-only, not idempotent, and not destructive. The description adds valuable context: it saves to local storage, may consume quota, and does not auto-retry on failure. This goes beyond annotation data without contradicting it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences: the first front-loads the core action and return values, the second covers model selection and important caveats about quota and retries. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description's mention of returned absolute path, actual dimensions, and PNG preview is essential and adequate. It also covers model default and failure behavior. Minor gaps remain, such as where the file is saved and overwrite behavior, but these are not critical for invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and two parameters already have schema descriptions. The description adds the default model (gpt-image-2.5-flare) and allowed alternatives, which is useful beyond the schema. However, it does not add meaning for the prompt parameter beyond what the schema implies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb and resource: generate an image via CLIProxy's images/generations endpoint, save it locally, and return absolute path, actual dimensions, and PNG preview. This clearly distinguishes it from sibling tools list_models and generate_text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or alternatives are mentioned. The description explains how the tool behaves (default model, quota, retry policy) but does not tell the agent when to choose this over generate_text or list_models.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_textA

通过 CLIProxy 生成文本。支持 Chat Completions 和 Responses;调用可能消耗额度。

ParametersJSON Schema
NameRequiredDescriptionDefault
apiNo
modelNoCLIProxy 模型 ID;省略时使用配置中的默认模型
promptYes
systemNo
maxTokensNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag that the tool is not read-only and not idempotent, so the mutation profile is known. The description adds the genuine trait '调用可能消耗额度' (may consume quota), which is a cost/side-effect caveat beyond the structured data, plus the API-mode support. This is useful but modest context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the purpose is front-loaded and the quota warning is appended efficiently. Every clause provides information that is not redundant with the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no output schema and 20% schema coverage, the description is too sparse. It omits parameter semantics, return format, error behavior, and any guidance on choosing api or system; an agent would have to guess or rely on external knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%—only the 'model' parameter is described in the schema. The description adds no explanation of prompt, system, api, or maxTokens, so it fails to compensate for the uncovered parameters. The required prompt parameter is completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with '通过 CLIProxy 生成文本'—a specific verb (generate) and resource (text via CLIProxy). It also names the two supported API variants (Chat Completions and Responses), which clearly distinguishes it from sibling tools list_models and generate_image by semantic domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use or when-not-to-use guidance. The sibling tool names imply the distinction (text vs. image generation vs. model listing), but there is no statement about which tool to select for which task or when to avoid this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA
Read-only

查询已配置 CLIProxy 的可用模型 ID。列表不保证所有模型支持所有接口。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark it read-only and open-world. The description adds value by specifying the list covers 'configured CLIProxy' models and by disclosing that a listed model may not support all interfaces. There is no contradiction with the readOnlyHint/openWorldHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise sentence states the core action and resource, followed by an important caveat. Every clause earns its place and the main purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter read-only listing tool, the description is largely sufficient: it names the return content (model IDs) and an important limitation. It could be more complete by explicitly connecting the output to the generate_text/generate_image siblings, but the low complexity keeps the gap minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema fully covers that with an empty properties object, so the baseline is 4. There are no parameter semantics for the description to enrich.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb '查询' (query) with resource '已配置 CLIProxy 的可用模型 ID', making it clear this tool lists configured model IDs. This distinguishes it from the sibling generation tools generate_text and generate_image, which produce content rather than enumerate models.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context implies list_models is used to discover model IDs before generation, but the description never explicitly says when to use it versus the siblings or that the IDs feed into generate_text/generate_image. The caveat about interface support hints at selection considerations, but it stops short of real usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.2.2
    • First observedgenerate_image
    • First observedgenerate_text
    • First observedlist_models

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: listing available models, generating text, and generating images. There is no overlap or ambiguity between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun snake_case pattern: list_models, generate_text, generate_image. The naming is predictable and easily understood.

Tool Count5/5

Three tools is well-scoped for a narrow proxy service focused on model discovery, text generation, and image generation. Each tool earns its place without unnecessary bloat.

Completeness5/5

The tool surface covers the core lifecycle of the domain: discovering available models, generating text, and generating images. No obvious critical operations are missing for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers