Cloudflare Workers AI MCP Server
# Cloudflare Workers AI MCP Server
Model Context Protocol (MCP) server that gives AI agents access to **Cloudflare Workers AI** — serverless LLM inference, embeddings, and image generation with a generous free tier.
## Tools
| Tool | Description |
|---|---|
| `list_models` | List supported chat, embedding, and image models |
| `chat_completion` | LLM chat completion (Llama 3.3 70B, Llama 3.1 8B, Llama 4 Scout, Qwen Coder 32B, DeepSeek R1 Distill) |
| `embed_text` | Text embeddings (BGE small/base) |
| `generate_image` | Image generation (Flux 1 Schnell) → base64 image (PNG/JPEG) |
## Setup
1. Create a Cloudflare API token with the **Workers AI** permission:
https://dash.cloudflare.com/profile/api-tokens
2. Get your **Account ID** (right sidebar of the Cloudflare dashboard, or the `/accounts/{id}` segment of any dashboard URL)
3. Export the env vars:
```bash
export CLOUDFLARE_ACCOUNT_ID="your-account-id"
export CLOUDFLARE_API_TOKEN="your-api-token"
```
## Run
```bash
npm install && npm run build
npm start # stdio MCP server
```
### Claude Desktop
Add to `claude_desktop_config.json`:
```json
{
"mcpServers": {
"cloudflare-workers-ai": {
"command": "node",
"args": ["/absolute/path/to/cloudflare-workers-ai-mcp/dist/index.js"],
"env": {
"CLOUDFLARE_ACCOUNT_ID": "your-account-id",
"CLOUDFLARE_API_TOKEN": "your-api-token"
}
}
}
}
```
### Cursor / VS Code / other MCP clients
Point the client at the same command (`node dist/index.js` with the two env vars).
### OpenClaw
OpenClaw supports MCP servers — add the same command to its MCP configuration.
## Example prompts
- "Summarize this text using the Llama 3.3 70B model on Cloudflare Workers AI"
- "Generate an image of a red fox in a snowstorm"
- "Embed these 3 sentences for a similarity search"
## Models (verified live)
Chat: `POST /ai/v1/chat/completions` (OpenAI-compatible) · Embeddings & images: `POST /ai/run/{model}` (native)
Models — chat: `@cf/meta/llama-3.3-70b-instruct-fp8-fast` · `@cf/meta/infire-llama-3.1-8b-instruct` · `@cf/meta/llama-4-scout-17b-16e-instruct` · `@cf/qwen/qwen2.5-coder-32b-instruct` · `@cf/deepseek-ai/deepseek-r1-distill-qwen-32b`
Embeddings: `@cf/baai/bge-small-en-v1.5` · `@cf/baai/bge-base-en-v1.5`
Images: `@cf/black-forest-labs/flux-1-schnell`
> Note: `@cf/meta/llama-3.1-8b-instruct` was deprecated by Cloudflare (2026-05-30) — use `@cf/meta/infire-llama-3.1-8b-instruct`.
## Test
```bash
npm test # unit tests (mocked fetch) + protocol test
```
## License
MIT — built by [puspoaditya](https://github.com/puspoaditya).
TDQS
Scored across 4 tools
Each tool addresses a distinct capability: discovery (list_models), text generation (chat_completion), vector embeddings (embed_text), and image generation (generate_image). There is no overlap in purpose, making selection unambiguous.
All tool names follow a clear verb_noun pattern (list_models, chat_completion, embed_text, generate_image) with consistent snake_case. The style is uniform and predictable, aiding agent understanding.
With only four tools, the server is tightly scoped to core AI inference tasks (listing, chat, embeddings, images). Each tool is essential and the count is well below the threshold for bloat, making the surface easy to navigate.
The server covers the three primary inference modalities advertised (chat, embeddings, images) plus model discovery. Minor gaps exist (e.g., audio or translation tasks), but for its stated purpose as a Workers AI inference wrapper, the surface is reasonably complete.