Skip to main content
Glama
README.md
# glama-gateway-mcp

A Models Context Protocol (MCP) server that exposes the **Glama AI gateway**
(https://gateway.glama.ai/v1, an OpenAI-compatible endpoint for 100+ models)
as MCP tools. Standard library only, Python 3.9+, no model downloads.

## Statement of need

MCP clients (Claude Desktop, Cursor, and a growing list of agent runtimes)
each expect an MCP server per capability. When a research project needs to
call language models across many providers — `openai/...`, `anthropic/...`,
`google/...`, and dozens more — wiring every client to every provider is a
tangle of credentials and bespoke plugins. Glama centralizes provider access
behind one OpenAI-compatible API, but it does not, by itself, hand a model
to an MCP client.

`glama-gateway-mcp` closes exactly that seam. It is a thin, dependency-free
MCP server that fronts the gateway: an agent (or IDE, or personal assistant)
talks MCP to this one server and gains `openai/gpt-4o`, `anthropic/claude-2`,
and any other gateway model through four tools — listing models, chat
completion, streaming completion, and request-status lookup. One API key, one
credential, one endpoint for every connected application.

Distinguishing design choices:

- **stdio MCP, implemented on the stdlib** — the protocol layer uses only
  `json`, `urllib` and `sys`, so it runs anywhere Python runs and is trivial
  to audit.
- **OpenAI-compatible request/response shapes** — the gateway returns raw
  gateway bodies, so no field mapping is ever wrong.
- **Server-side stream reassembly** — MCP clients that cannot hold a live SSE
  stream still receive the full completion text plus usage metadata.
- **Clean JSON-RPC errors** — a missing key, unreachable gateway, or unknown
  tool yields a readable MCP error instead of a hang or silent failure.

## Install

```sh
pip install .
# development:
pip install -e ".[dev]"
```

Requires Python 3.9+. Runtime dependencies: none.

## Usage

Export your Glama key, then run the server:

```sh
export GLAMA_API_KEY="your_key_here"
export GLAMA_DEFAULT_MODEL="openai/gpt-4o"    # optional
glama-mcp
```

or run it as a module:

```sh
python -m glama_mcp
```

Connect any MCP client to the `glama-mcp` stdio command. The exposed tools:

| Tool | Purpose |
|---|---|
| `glama_list_models` | list models available through the gateway |
| `glama_chat_completion` | one-shot chat completion |
| `glama_stream_completion` | streamed completion (reassembled) |
| `glama_request_status` | status of a completion request by id |

### Per-application wiring

Ready-to-paste `mcpServers` blocks for individual applications live in
[`configs/`](configs/): Celebrum, Samvit, Collabuild, S-AI, hermes-agent,
teddy-techlearn, health-quest, ai-content-studio, and the portfolio site.
Each block registers this server for that application's MCP client.

```json
{
  "mcpServers": {
    "glama-gateway": {
      "command": "glama-mcp",
      "env": { "GLAMA_API_KEY": "your_key_here" }
    }
  }
}
```

### Verifying the server by hand

```sh
printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{}}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | python -m glama_mcp
```

## Tests

```sh
python -m pytest
```

The suite covers the gateway client (payload shapes, SSE reassembly), protocol
handshake, tool discovery and dispatch, error handling, missing-key behaviour,
and a full stdio round trip.

## Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md). Keep it dependency-free and Python 3.9
compatible; every change needs a test.

## JOSS paper

Submission materials in [`paper/`](paper/): `paper.md`, `paper.bib`, and
`JOSS_SUBMISSION_READINESS.md` (records which gates are met and which still
require calendar time).

## License

MIT — see [LICENSE](LICENSE).

## AI usage disclosure

Code and documentation were drafted with generative-AI assistance and reviewed
by the human maintainer, who made the design decisions. This disclosure is kept
in line with the JOSS AI usage policy.

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation4/5

list_models and request_status are clearly distinct, but chat_completion and stream_completion overlap in purpose as they both perform chat completions. The descriptions clarify streaming vs. non-streaming, so an agent can differentiate them, though some ambiguity remains.

Naming Consistency5/5

All tool names share the glama_ prefix and follow a consistent snake_case verb_noun pattern. The naming is predictable and uniform across the set.

Tool Count5/5

Four tools is well-scoped for a gateway-focused server covering model discovery, completion, streaming, and status lookup. Each tool has a clear, non-redundant role.

Completeness4/5

The core workflow of listing models and running both standard and streaming completions is covered, plus status lookup for async requests. Minor additions like request cancellation or model details would improve completeness, but no critical gap exists.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive