Agnes AI MCP Server
ο»Ώ<div align="center">
# π¨ Agnes AI MCP Server
**Free Text-to-Image & Text-to-Video generation via [Agnes AI](https://agnes-ai.com)**
[](https://pypi.org/project/agnes-mcp/)
[](https://pypi.org/project/agnes-mcp/)
[](https://github.com/MSWEIMZ/agnes-mcp/actions/workflows/ci.yml)
[](https://opensource.org/licenses/MIT)
[](https://www.python.org/downloads/)
[](https://modelcontextprotocol.io)
English | [δΈζ](README_CN.md)
</div>
---
## π Quick Start
```bash
# 1. Install (one command)
pip install agnes-mcp
# 2. Get a free API key at https://agnes-ai.com
# 3. Add to your MCP client config:
```
**Claude Desktop / Cursor / Windsurf** (`claude_desktop_config.json` or equivalent):
```json
{
"mcpServers": {
"agnes-mcp": {
"command": "uvx",
"args": ["agnes-mcp"],
"env": {
"AGNES_API_KEY": "your-api-key-here"
}
}
}
}
```
**Codex** (`config.toml`):
```toml
[mcp_servers.agnes_mcp]
command = "uvx"
args = ["agnes-mcp"]
[mcp_servers.agnes_mcp.env]
AGNES_API_KEY = "your-api-key-here"
```
That's it! Now you can generate images and videos directly from your AI assistant.
---
## β¨ Why Agnes MCP?
| Feature | Agnes MCP | Other AI Image Services |
|---------|-----------|------------------------|
| **Price** | **$0 / image, $0 / second** | $0.02 - $0.08 / image |
| Text-to-Image | β
2 models (2.0 & 2.1 Flash) | β
Usually 1 model |
| Image-to-Image | β
Reference image + prompt | β or limited |
| Batch Generation | β
1-4 images at once | β |
| Text-to-Video | β
Up to 18s, 1080p | β or paid only |
| Image-to-Video | β
Static image β video | β or paid only |
| Multi-image Video | β
Keyframe animation | β |
| Auto Download | β
Saves locally automatically | β Manual download |
| MCP Standard | β
Full compliance | Varies |
**Yes, it's completely free.** Agnes AI currently offers all image and video generation at $0. Just register and get an API key.
---
## πΌοΈ Demo
### Text-to-Image (agnes-image-2.1-flash)
> *"A majestic dragon flying over a Chinese mountain landscape at sunset, cinematic lighting, epic fantasy art"*

### Text-to-Image (agnes-image-2.0-flash)
> *"A cozy Japanese ramen shop at night, warm lantern light, rain falling, anime style"*

---
## π¦ Tools
| Tool | Description | Example |
|------|-------------|---------|
| `text_to_image` | Generate image(s) from text | `prompt: "a cat"` + optional `n: 4`, `images: [ref_url]` |
| `image_to_image` | Generate from reference image(s) + text | `prompt: "make it cyberpunk"` + `images: [url]` |
| `text_to_video` | Generate video from text/image(s) | `prompt: "a cat dancing"` + optional `mode`, `num_inference_steps` |
| `image_to_video` | Animate a static image into video | `prompt: "zoom in slowly"` + `image: "url"` |
| `keyframe_animation` | Smooth transition between keyframe images | `prompt: "morph scene"` + `images: [url1, url2, ...]` |
| `check_video_status` | Check async video task status | `video_id: "xxx"` or `task_id: "xxx"` |
---
## βοΈ Environment Variables
| Variable | Required | Default | Description |
|----------|----------|---------|-------------|
| `AGNES_API_KEY` | **Yes** | - | Your Agnes AI API key |
| `AGNES_API_BASE` | No | `https://apihub.agnes-ai.com/v1` | API base URL |
| `AGNES_DEFAULT_MODEL` | No | `agnes-image-2.1-flash` | Default image model |
| `AGNES_DEFAULT_SIZE` | No | `1024x768` | Default image size |
---
## π Get a Free API Key
1. Visit [https://agnes-ai.com](https://agnes-ai.com)
2. Create an account (free)
3. Go to Console β API Keys β Create
4. Copy the key and paste into your config
---
## β
Supported Clients
- [x] **Claude Desktop**
- [x] **Codex (OpenAI)**
- [x] **Cursor**
- [x] **Windsurf**
- [x] **Cherry Studio**
- [x] Any MCP client with `stdio` transport
---
## π Changelog
### v0.3.0 (2026-06-28)
- β¨ New tool: `image_to_video` β animate a static image into video
- β¨ New tool: `keyframe_animation` β smooth transitions between multiple keyframe images
- β¨ `text_to_video`: added `mode` and `num_inference_steps` parameters
- β¨ `create_video_task` / `generate_video`: support `mode` (e.g. `ti2vid`, `keyframes`) and `num_inference_steps`
- β
28 tests passing
### v0.2.0 (2026-06-27)
- β¨ New tool: `image_to_image` β generate from reference image(s) + prompt
- β¨ `text_to_image`: batch generation (`n: 1-4`) and multi-image composition (`images`)
- β¨ `text_to_video`: multi-image video / keyframe animation (`images`)
- π Unified multi-image download logic
- β
19 tests passing
### v0.1.1 (2026-06-26)
- π Initial public release
- text_to_image, text_to_video, check_video_status
- Async httpx with retry mechanism
- Auto-download to local filesystem
---
## π€ Contributing
See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.
---
## π License
MITTDQS
Scored across 6 tools
The primary tools are clearly separated by source and target modality (text/image/video), making their main purposes easy to distinguish. However, text_to_video also accepts optional image(s) and keyframe modes, which creates some boundary overlap with image_to_video and keyframe_animation.
Four tools follow a clean source_to_target naming pattern (text_to_image, image_to_image, text_to_video, image_to_video), and all names use snake_case. keyframe_animation and check_video_status break the pattern slightly, but the convention is still predictable and readable.
Six tools is a well-scoped size for a media generation server, covering both image and video generation without redundancy or bloat. Each tool provides a distinct high-level capability, and the count feels appropriate for the domain.
The set covers the core generation workflows: text-to-image, image-to-image, text-to-video, image-to-video, and keyframe animation, plus async status checking. Minor gaps exist, such as no model listing or task cancellation, but agents can complete all primary generation tasks without dead ends.