Skip to main content
Glama

OpenAI GPT MCP Server

A local Model Context Protocol server that provides OpenAI GPT text generation over standard input/output.

Prerequisites

  • Node.js 18 or newer

  • An OpenAI API key

Related MCP server: gpt-mcp

Install and build

npm install
npm run build

Copy .env.example to .env and set your API key:

OPENAI_API_KEY=your_api_key

OPENAI_MODEL is optional and defaults to gpt-5. OPENAI_IMAGE_MODEL is optional and defaults to gpt-image-2.

MCP client configuration

Build first, then add this server to your MCP client's configuration. Replace <absolute-path-to-repository> with the absolute path to your clone.

{
  "mcpServers": {
    "openai-gpt": {
      "command": "node",
      "args": ["<absolute-path-to-repository>\\dist\\index.js"]
    }
  }
}

Tool: generate_text

Required input:

  • prompt: text to send to the model.

Optional inputs:

  • instructions: system-level behavior for the response.

  • model: a GPT model ID, overriding OPENAI_MODEL.

  • max_output_tokens: maximum response size, from 1 through 16,384.

The server uses the OpenAI Responses API and returns generated text to the MCP client. It never writes API keys to output or logs.

Tool: generate_image

Generates one PNG image and returns it directly to the MCP client.

  • prompt (required): detailed image description.

  • model (optional): image model ID, overriding OPENAI_IMAGE_MODEL.

  • size (optional): 1024x1024, 1536x1024, or 1024x1536.

  • quality (optional): low, medium, or high.

Image generation consumes OpenAI API credits separately from Claude usage.

Run directly

npm start

For interactive inspection:

npm run inspect

Available Tools

2 tools
generate_imageGenerate an image with OpenAIB

Generate one image from a text prompt using OpenAI's image generation API.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoImage dimensions. Defaults to 1024x1024.
modelNoOptional image model ID. Defaults to gpt-image-2.
promptYesA detailed description of the image to generate.
qualityNoImage quality. Defaults to medium.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states that it generates one image, which is minimal. It does not mention output format (e.g., URL or file), potential costs, rate limits, or any side effects. For a generation tool that might involve external API calls and delays, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is front-loaded with the core action. It is concise and free of filler or redundancy, earning full marks for efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (4 parameters, 2 enums, no output schema), the description is under-specified. It covers the basic action but lacks usage guidance, behavioral notes, and any explanation of what the response will contain. While parameters are well-documented in schema, the overall description is too sparse for a complete evaluation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already describes all parameters. The description adds no additional meaning beyond restating that it uses a text prompt, which is already in the prompt parameter's description. Per rubric, baseline is 3 for high schema coverage, and no extra value is added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: generating one image from a text prompt using OpenAI's image generation API. It specifies the verb 'generate', the resource 'image', and distinguishes from the sibling generate_text by focusing on image output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The sibling tool generate_text is not mentioned, and there are no conditions, exclusions, or context for selecting image generation over text generation. The description provides only the basic action without usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_textGenerate text with OpenAIA

Generate a text response using an OpenAI GPT model via the Responses API.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOptional model ID. Defaults to gpt-5.
promptYesThe user prompt to send to the model.
instructionsNoOptional system-level instructions for the model.
max_output_tokensNoOptional maximum number of generated tokens (up to 16,384).

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It mentions the API used ('OpenAI GPT model via the Responses API') which hints at behavior but does not disclose details like rate limits, costs, or that the tool invokes an external service. It does not state what happens with token limits or error handling. Given the lack of annotations, this is a minimal but adequate level of behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence, front-loaded with the purpose, and contains no fluff. It is appropriately concise and focused.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a medium-complexity tool with 4 parameters, all documented in the schema, and no output schema. The description covers the primary action and API, but lacks details about return structure (text response) and any constraints like model availability or fallbacks. However, given schema completeness and simplicity, a score of 3 is reasonable—adequate but not rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no additional parameter-specific details beyond what the schema already states. It doesn't explain the implications of max_output_tokens or instructions, but schema descriptions are sufficient. No contradictions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (generate) and the resource (text response using an OpenAI GPT model via the Responses API). It distinguishes from sibling generate_image by specifying 'text response' and 'GPT model', though it does not explicitly contrast with that tool. The title and description are aligned and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for generating text from a prompt but does not explicitly specify when to use it vs. alternatives. With only one sibling (generate_image), context strongly suggests text generation, but no explicit when/when-not guidance is given. The description could mention that this is for textual completions as opposed to image generation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.0
    • First observedgenerate_image
    • First observedgenerate_text

TDQS

A3.7/5.0

Scored across 2 tools

Disambiguation5/5

The two tools are clearly distinct: one generates text and the other generates images. There is no overlap or ambiguity in their purposes.

Naming Consistency5/5

Both tools follow the same consistent generate_noun pattern using snake_case. The naming convention is uniform and predictable.

Tool Count3/5

Two tools is borderline for a coherent server; they are well-scoped but the surface feels thin for a general OpenAI GPT server. Each tool earns its place, but the count is on the low end.

Completeness4/5

The two-generation capabilities cover the core domain of text and image generation with no obvious dead ends. Minor gaps exist if considering broader OpenAI features like embeddings or audio, but for the stated GPT-focused purpose the coverage is reasonable.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers