apimodels-mcp
# apimodels-mcp
MCP server for [apimodels.app](https://apimodels.app) — call **image, video, LLM chat and text-to-speech** models with one API key, from Claude Desktop, Cursor, or any MCP client.
One key unlocks GPT-5.5, Claude, Gemini, GLM, DeepSeek, Qwen, Seedance, Veo, Kling, gpt-image-2, Gemini Image, MiniMax speech and more — billed in USD, you only pay for successful generations.
## Tools
| Tool | What it does |
|------|--------------|
| `list_models` | List available model ids (chat / image / video / audio). |
| `chat` | Chat / text completion with any LLM (`gpt-5-5`, `claude-opus-4-8`, `gemini-3-pro-preview`, …). |
| `generate_image` | Text-to-image or image edit; returns the image URL(s) plus a downscaled preview the model can look at. With `doubao-seedream-5-0-flash` it can also keep a transparent background (`background: "transparent"`) or split one image into a base plus up to 16 transparent layers with names and positions (`layer_decomposition: true`, billed per output image). |
| `review_image` | A vision model critiques an image against your brief and proposes a revised prompt. |
| `generate_video` | Text-to-video (optional reference image); returns the video URL(s), or a task id if it is not done within `wait_seconds`. |
| `get_task` | Wait for / check on a task that `generate_image`, `generate_video` or `text_to_speech` handed back as still running. |
| `text_to_speech` | Text-to-speech (MiniMax voices); returns the audio URL. ElevenLabs TTS is not exposed here — it streams raw bytes from `POST /v1/tts/stream` rather than returning a URL. |
### Long generations do not get lost
Every generation is asynchronous on apimodels, and video is slow: a median of about 2.5 minutes, 9 in 10 within 8 minutes (production, week to 2026-09-22). Images take about 50 seconds. Meanwhile Codex aborts an MCP tool call after 60 seconds by default, and the MCP SDK's own client timeout is 60 seconds too. A tool that blocks until the video is ready therefore gets killed mid-wait — the task keeps running, the account is billed when it finishes, and the assistant never sees the URL. Versions up to 0.2.x did exactly that.
Since 0.3.0 the generation tools wait at most `wait_seconds` (default 50) and then return the task id with a "still running" note; the assistant calls `get_task`, which waits up to another `wait_seconds` and returns the URL(s) — with the image preview for image tasks — or "still running" again. Nothing is resubmitted and nothing is billed twice. The assistant does this on its own; you just ask for the video.
- **Codex**: keep the default. Or raise `tool_timeout_sec` for this server in `config.toml` and pass a larger `wait_seconds`.
- **Claude Code / Claude Desktop / Cursor**: no 60-second limit, so `wait_seconds: 600` on `generate_video` gets the URL in one call. `APIMODELS_WAIT_SECONDS=600` in the server's env makes that the default.
### The model can check its own work
Ask for an image and let the assistant iterate until it is right — "make a 16:9 banner that says SAVE 10%, check the spelling, fix it if needed":
1. `generate_image` returns the URL **and a preview of the image itself** (max 1024px JPEG). Clients that pass tool-result images to the model — Claude Desktop, Claude Code, Cursor — let it see what it made. Pass `return_image: false` to skip the preview.
2. `review_image` works everywhere, including clients that show tool-result images to you but not to the model (Cherry Studio is one). It sends the image and your brief to a vision model and returns what matches, what is wrong (garbled text, composition, aspect ratio, artifacts) and a revised prompt. One review costs well under $0.01 on the default `gpt-5.6-luna`.
The assistant picks `aspect_ratio` and `resolution` itself from what you ask for, so "make it 16:9" in plain words is enough.
### Local images just work
`image_url` on `generate_image` and `generate_video` takes any of these:
- a public `https://…` URL — passed through untouched
- **a local file path** — `/Users/me/photo.png`, `./ref.jpg`, `~/Pictures/x.webp`
- **a URL on your own machine** — `http://127.0.0.1:8000/photo.png`, `http://localhost:3000/…`
- a `data:image/png;base64,…` URI
The last three are uploaded for you first, and the resulting public URL is what gets
generated from. This has to happen here rather than server-side: the file exists only on
your machine, and `127.0.0.1` means *our* server when our server resolves it — which is why
passing one to the REST API directly fails with `private/reserved IP addresses not allowed`.
This MCP server runs next to your files, so it can do what our servers cannot.
Uploads land in your account's R2 space and are auto-deleted after 7 days.
## Setup
1. Get an API key at <https://apimodels.app/console/api-keys> (it looks like `sk_…`).
2. Add the server to your MCP client.
### Claude Desktop
Edit `claude_desktop_config.json` (Settings → Developer → Edit Config):
```json
{
"mcpServers": {
"apimodels": {
"command": "npx",
"args": ["-y", "apimodels-mcp"],
"env": {
"APIMODELS_API_KEY": "sk_your_key_here"
}
}
}
}
```
Restart Claude Desktop. You can now ask it to "generate an image of …" or "make a 5-second video of …".
### Cursor
`Settings → MCP → Add new MCP server`, or add to `~/.cursor/mcp.json`:
```json
{
"mcpServers": {
"apimodels": {
"command": "npx",
"args": ["-y", "apimodels-mcp"],
"env": { "APIMODELS_API_KEY": "sk_your_key_here" }
}
}
}
```
### Codex CLI
Add to `~/.codex/config.toml`:
```toml
[mcp_servers.apimodels]
command = "npx"
args = ["-y", "apimodels-mcp"]
env = { APIMODELS_API_KEY = "sk_your_key_here" }
# Optional. Codex aborts a tool call after 60 s by default; the tools stay under that
# on their own (see "Long generations do not get lost"), so this is only needed if you
# want generate_video to return the URL in one call — then also pass wait_seconds: 600.
# tool_timeout_sec = 660
```
### Cherry Studio
In `Settings → MCP Servers`, add a new server of type **stdio**:
- Command: `npx`
- Arguments: `-y apimodels-mcp`
- Environment variables: `APIMODELS_API_KEY=sk_your_key_here`
Enable the server, then select it for your conversation from the MCP control under the chat box. Use a chat model that supports tool calls (Claude, GPT, Gemini …) as the conversation model — it calls the image model for you. Cherry Studio needs Node.js installed for `npx`; on Windows install it from <https://nodejs.org>.
Any other MCP client works the same way — run `npx -y apimodels-mcp` over stdio with `APIMODELS_API_KEY` in the environment.
## Models, docs and pricing
Everything the tools call is documented on apimodels.app:
- [API documentation](https://apimodels.app/docs) · [pricing](https://apimodels.app/pricing) · [full model catalog](https://apimodels.app/models)
- Image: [GPT Image 2.5 API](https://apimodels.app/docs/gpt-image-2-5) ([model page](https://apimodels.app/models/gpt-image-2.5-flare)), [GPT Image 2 API](https://apimodels.app/docs/gpt-image-2), [Nano Banana 2.1 API](https://apimodels.app/docs/nano-banana-2-1) ([model page](https://apimodels.app/models/nano-banana-2-1)), [all image models](https://apimodels.app/docs/image)
- Video: [Seedance 2.5 API](https://apimodels.app/docs/seedance-2-5), [Google Veo API](https://apimodels.app/docs/google-veo), [MiniMax H3 API](https://apimodels.app/docs/minimax-h3), [FlashVSR video upscaling](https://apimodels.app/docs/flashvsr)
- Chat and speech: [LLM API (GPT, Claude, Gemini, DeepSeek, GLM, Qwen)](https://apimodels.app/docs/llm), [Claude Sonnet 5.5](https://apimodels.app/models/claude-sonnet-5-5), [audio and text-to-speech](https://apimodels.app/docs/audio)
- Other ways in: [Claude Code setup](https://apimodels.app/docs/claude-code), [chat clients](https://apimodels.app/docs/clients), [Agent Skills](https://apimodels.app/docs/skills), [free calculators and tools](https://apimodels.app/tools)
- Prompt libraries with example outputs: [GPT Image 2.5 prompts](https://apimodels.app/gpt-image-2-5-prompts), [GPT Image 2 prompts](https://apimodels.app/gpt-image-2-prompts), [Seedance 2.5 prompts](https://apimodels.app/seedance-2-5-prompts), [MiniMax H3 prompts](https://apimodels.app/minimax-h3-prompts)
## Configuration
| Env var | Default | Description |
|---------|---------|-------------|
| `APIMODELS_API_KEY` | — (required) | Your `sk_…` key. |
| `APIMODELS_BASE_URL` | `https://api.apimodels.app/v1` | API base URL. |
| `APIMODELS_WAIT_SECONDS` | `50` | Default `wait_seconds` for `generate_image`, `generate_video`, `text_to_speech` and `get_task`: how long a call waits before handing back a task id. Max 900. |
| `APIMODELS_TIMEOUT_MS` | — | Deprecated (0.2.x): the same wait in milliseconds. Still honoured if set. |
## Local development
```bash
pnpm install
pnpm build
APIMODELS_API_KEY=sk_... node dist/index.js # runs over stdio
```
## License
MIT
TDQS
Scored across 6 tools
Each tool maps to a clearly distinct modality or action: model listing, text chat, image generation, image critique, video generation, and speech synthesis. There is no meaningful overlap between tool purposes, so an agent should be able to select correctly.
Most tools follow a snake_case verb_noun pattern (list_models, generate_image, review_image, generate_video). text_to_speech is a descriptive noun phrase and chat is a single verb, which are minor deviations from the dominant pattern.
Six tools is a well-scoped set for a multimodal model gateway covering chat, image, video, and audio. Each tool is independently useful and there are no redundant additions.
The tool surface covers the main generation workflows for all four advertised modalities and includes a helpful image-review loop. Minor gaps exist, such as no video review or audio transcription, but agents can complete core tasks without dead ends.