Skip to main content
Glama
puspoaditya

Cloudflare Workers AI MCP Server

README.md
# Cloudflare Workers AI MCP Server

Model Context Protocol (MCP) server that gives AI agents access to **Cloudflare Workers AI** — serverless LLM inference, embeddings, and image generation with a generous free tier.

## Tools

| Tool | Description |
|---|---|
| `list_models` | List supported chat, embedding, and image models |
| `chat_completion` | LLM chat completion (Llama 3.3 70B, Llama 3.1 8B, Llama 4 Scout, Qwen Coder 32B, DeepSeek R1 Distill) |
| `embed_text` | Text embeddings (BGE small/base) |
| `generate_image` | Image generation (Flux 1 Schnell) → base64 image (PNG/JPEG) |

## Setup

1. Create a Cloudflare API token with the **Workers AI** permission:
   https://dash.cloudflare.com/profile/api-tokens
2. Get your **Account ID** (right sidebar of the Cloudflare dashboard, or the `/accounts/{id}` segment of any dashboard URL)
3. Export the env vars:

```bash
export CLOUDFLARE_ACCOUNT_ID="your-account-id"
export CLOUDFLARE_API_TOKEN="your-api-token"
```

## Run

```bash
npm install && npm run build
npm start   # stdio MCP server
```

### Claude Desktop

Add to `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "cloudflare-workers-ai": {
      "command": "node",
      "args": ["/absolute/path/to/cloudflare-workers-ai-mcp/dist/index.js"],
      "env": {
        "CLOUDFLARE_ACCOUNT_ID": "your-account-id",
        "CLOUDFLARE_API_TOKEN": "your-api-token"
      }
    }
  }
}
```

### Cursor / VS Code / other MCP clients

Point the client at the same command (`node dist/index.js` with the two env vars).

### OpenClaw

OpenClaw supports MCP servers — add the same command to its MCP configuration.

## Example prompts

- "Summarize this text using the Llama 3.3 70B model on Cloudflare Workers AI"
- "Generate an image of a red fox in a snowstorm"
- "Embed these 3 sentences for a similarity search"

## Models (verified live)

Chat: `POST /ai/v1/chat/completions` (OpenAI-compatible) · Embeddings & images: `POST /ai/run/{model}` (native)

Models — chat: `@cf/meta/llama-3.3-70b-instruct-fp8-fast` · `@cf/meta/infire-llama-3.1-8b-instruct` · `@cf/meta/llama-4-scout-17b-16e-instruct` · `@cf/qwen/qwen2.5-coder-32b-instruct` · `@cf/deepseek-ai/deepseek-r1-distill-qwen-32b`

Embeddings: `@cf/baai/bge-small-en-v1.5` · `@cf/baai/bge-base-en-v1.5`

Images: `@cf/black-forest-labs/flux-1-schnell`

> Note: `@cf/meta/llama-3.1-8b-instruct` was deprecated by Cloudflare (2026-05-30) — use `@cf/meta/infire-llama-3.1-8b-instruct`.

## Test

```bash
npm test   # unit tests (mocked fetch) + protocol test
```

## License

MIT — built by [puspoaditya](https://github.com/puspoaditya).

TDQS

A4.7/5.0

Scored across 4 tools

Disambiguation5/5

Each tool addresses a distinct capability: discovery (list_models), text generation (chat_completion), vector embeddings (embed_text), and image generation (generate_image). There is no overlap in purpose, making selection unambiguous.

Naming Consistency5/5

All tool names follow a clear verb_noun pattern (list_models, chat_completion, embed_text, generate_image) with consistent snake_case. The style is uniform and predictable, aiding agent understanding.

Tool Count5/5

With only four tools, the server is tightly scoped to core AI inference tasks (listing, chat, embeddings, images). Each tool is essential and the count is well below the threshold for bloat, making the surface easy to navigate.

Completeness4/5

The server covers the three primary inference modalities advertised (chat, embeddings, images) plus model discovery. Minor gaps exist (e.g., audio or translation tasks), but for its stated purpose as a Workers AI inference wrapper, the surface is reasonably complete.

Maintenance

ActivitySlowing
ResponsivenessNo issues