Skip to main content
Glama
remiehneppo

Code Generator MCP Server

by remiehneppo

Code Generator MCP Template Server

An MCP (Model Context Protocol) server built using Python's FastMCP framework. It exposes 4 precise, structured code-generation tools backed by prompt templates to generate production-grade, parsed code using local or cloud-based OpenAI-compatible APIs (such as llama-server, qwen-coder, coder-expert, or OpenAI's API).

๐Ÿš€ Features

Exposes 4 key MCP tools that use structured templates to instruct the model to think step-by-step and produce clean, executable code:

  1. generate_standard_function (Template 1): Generates a standalone function based on constraints, edge cases, test cases, and external integration notes.

  2. generate_codebase_context (Template 2): Generates a function that respects and integrates with existing codebase structures and dependencies.

  3. generate_bugfix_refactor (Template 3): Focuses on refactoring or repairing current buggy implementations based on problem descriptions and test expectations.

  4. generate_multi_function_module (Template 4): Generates a multi-function module, validating that there are no circular dependencies or undefined functions.

โš™๏ธ Parser & Guardrails

  • Markdown Stripper: Automatic code parsing (extract_code_from_response) strips any markdown code blocks (```python) generated by instruct models, guaranteeing only raw executable code is returned.

  • Reasoning Fallback: Correctly handles DeepSeek-style reasoning models or llama-server configurations where all output is redirected into the reasoning_content field instead of content.

  • Max Tokens Guardrail: Enforces a 2048 token limit per request to prevent local model reasoning loops and timeouts.


Related MCP server: mcp_server_for_claudes_toolbox

๐Ÿ› ๏ธ Configuration

Configure the server using command-line arguments or environment variables:

Setting

CLI Argument

Environment Variable

Default Value

Description

API URL

--api-url

CODE_GEN_API_URL / OPENAI_BASE_URL

https://api.openai.com/v1

OpenAI-compatible endpoint

Model

--model

CODE_GEN_MODEL / OPENAI_MODEL

gpt-4o

The model name to target

API Key

--api-key

CODE_GEN_API_KEY / OPENAI_API_KEY

(Empty)

API token (optional for local endpoints)

WARNING

For security reasons, do not pass--api-key via command-line arguments as it will be visible in plain text in the host process table. Use the environment variables instead.


๐Ÿ“ฆ Installation & Setup

You can install and configure the server either automatically using the installation script or manually via Python.

The project includes an install.sh script that automatically builds a standalone executable using PyInstaller, installs it to ~/.local/bin/code-generator-mcp, and configures your target AI coding agent.

  1. Run the script and specify your target agent:

    chmod +x install.sh
    ./install.sh <agent_type>

    Supported <agent_type> values:

    • claude-desktop (Claude Desktop global configuration)

    • claude-code (Claude CLI global configuration at ~/.claude.json)

    • cursor (Cursor editor global config at ~/.cursor/mcp.json)

    • codex (Codex agent global config at ~/.codex/config.toml)

    • github-copilot (VS Code workspace-specific config at .vscode/mcp.json)

    • windsurf (Windsurf IDE configuration at ~/.codeium/windsurf/mcp_config.json)

    • zed (Zed editor config at ~/.config/zed/settings.json)

    • agy (Antigravity settings config at ~/.gemini/settings.json)

  2. (Optional) Customize the endpoint and model in the agent's configuration file or environment variables after installation.


Option 2: Manual Setup via Python

Prerequisites

  • Python 3.10+

  • Dependencies installed in virtual environment:

    python -m venv .venv
    source .venv/bin/activate
    pip install -r requirements.txt

Running Locally

To run the MCP server directly over standard input/output (stdio):

python src/code_generator_mcp/server.py --api-url http://localhost:8008/v1 --model coder-expert

Testing during development

You can use mcp dev (from MCP CLI) to test the server interactively in a development UI:

mcp dev src/code_generator_mcp/server.py -- --api-url http://localhost:8008/v1 --model coder-expert

๐Ÿ”Œ Integration Setup

To manually use this server with your favorite MCP client (like Claude Desktop or Cursor):

Claude Desktop Configuration

Open your Claude Desktop config file (usually located at ~/.config/Claude/claude_desktop_config.json on Linux/macOS or %APPDATA%\Claude\claude_desktop_config.json on Windows) and add the following entry:

{
  "mcpServers": {
    "code-generator-mcp": {
      "command": "/path/to/project/.venv/bin/python",
      "args": [
        "/path/to/project/src/code_generator_mcp/server.py"
      ],
      "env": {
        "CODE_GEN_API_URL": "http://localhost:8008/v1",
        "CODE_GEN_MODEL": "coder-expert"
      }
    }
  }
}

๐Ÿงช Testing

The codebase includes a fully-featured unit and integration test suite using pytest. Run tests with:

.venv/bin/pytest

Available Tools

5 tools
generate_bugfix_refactorC

Generate fixed or refactored code (Template 3) using coder expert model.

If generate_test_file is True, it will co-generate a matching unit test suite.

IMPORTANT: All parameters (such as 'task', 'problem', 'expected_behavior', 'constraints', etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNo
problemYes
languageYes
test_casesYes
constraintsNo
current_codeYes
expected_behaviorYes
generate_test_fileNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It usefully discloses one real behavioral constraint (all text parameters must be supplied in English, with a translation instruction) and the test co-generation behavior, but it says nothing about permissions, output format, limits, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose in the first sentence, followed by two focused notes. The English/translation block is somewhat verbose but carries genuinely actionable information, so it earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter generation tool with zero schema descriptions, no annotations, and only partial parameter coverage, the description is thin. The output schema existing means return values need no explanation, but the input semantics and usage routing are largely missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all 8 parameters, so the description must compensate but does not. It names a few fields (task, problem, expected_behavior, constraints, generate_test_file) only to impose the English requirement, without explaining what language, current_code, or test_cases mean or how task differs from problem.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate fixed or refactored code') and references 'Template 3' plus the coder expert model, so the agent understands the operation. However, it does not distinguish itself from siblings like generate_standard_function or generate_multi_function_module, leaving the boundary inferred.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus the sibling generators (standard function, multi-function module). The only conditional behavior mentioned is generate_test_file co-generation, which is a parameter effect rather than a when-to-use rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_codebase_contextC

Generate code for a function with codebase context (Template 2) using coder expert model.

If generate_test_file is True, it will co-generate a matching unit test suite.

IMPORTANT: All parameters (such as 'task', 'description', 'constraints', etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
languageYes
signatureYes
test_casesYes
constraintsNo
descriptionYes
existing_codeYes
interacts_withNo
generate_test_fileNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose that setting generate_test_file co-generates a unit test suite and that a 'coder expert model' is used, plus a notable English-only input constraint. It says nothing about how existing_code is used, permissions, latency, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core capability is front-loaded in the first sentence, with the conditional test-generation behavior immediately after. The English-only block is verbose but earns its place as a hard invocation constraint; only the unexplained 'Template 2' reference is dead weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, 6-required generation tool with zero schema coverage and no annotations, the description is far too thin. An output schema exists so return values need not be explained, but the caller is given no way to know what signature, existing_code, interacts_with, or test_cases should contain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the schema contributes no meaning and the description must compensate. It only names 'task', 'description', and 'constraints' as examples in the English-language note and explains generate_test_file; signature, existing_code, interacts_with, test_cases, and language are left entirely undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate code for a function with codebase context'), so an agent knows this produces code from existing code context. However, it does not differentiate itself from the close sibling generate_standard_function, and 'Template 2' is opaque internal jargon that carries no meaning for the caller.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to choose this over generate_standard_function, generate_multi_function_module, or generate_bugfix_refactor, despite these being obvious alternatives. The only conditional behavior mentioned is the generate_test_file flag, which is not selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_multi_function_moduleB

Generate code for a multi-function module (Template 4) using coder expert model.

If generate_test_file is True, it will co-generate a matching unit test suite.

IMPORTANT: All parameters (such as 'task', 'module_purpose', 'context', 'functions' list with its descriptions/constraints, etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
contextNo
languageYes
functionsYes
test_casesYes
dependenciesNo
shared_typesNo
export_formatNo
module_purposeYes
generate_test_fileNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully discloses the coder expert model, optional co-generated unit test suite, and mandatory English parameter translation, but omits permissions, side effects, determinism, and other operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then states the conditional test-generation behavior, then the critical language requirement. Every sentence adds information and there is no redundant restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter code-generation tool with no annotations, the description is incomplete. It does not explain key parameters, usage relative to siblings, or behavioral constraints beyond English translation and optional test generation, even though output-schema existence means return values need not be described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema description coverage is 0% across 10 parameters, so the description must compensate. It only names or hints at task, module_purpose, context, functions, and generate_test_file, while language, test_cases, dependencies, shared_types, export_format, and most function-spec fields remain unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: generate code for a multi-function module (Template 4) using a coder expert model. The phrase 'multi-function module' distinguishes it from generate_standard_function, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no guidance on when to choose this over generate_standard_function, generate_codebase_context, or generate_bugfix_refactor. The only conditional behavior mentioned is if generate_test_file is True, which is execution behavior rather than tool selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_standard_functionC

Generate code for a standard function spec (Template 1) using coder expert model.

If generate_test_file is True, it will co-generate a matching unit test suite.

IMPORTANT: All parameters (such as 'task', 'description', 'constraints', 'edge_cases', etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
contextNo
languageYes
signatureYes
edge_casesNo
test_casesYes
constraintsNo
descriptionYes
integration_noteNo
generate_test_fileNo
dependencies_allowedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that a coder expert model is used, that a test suite is co-generated when generate_test_file is true, and that all parameter values must be English. However it omits mutation/return behavior and error conditions, so the behavioral profile is only partially covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with no filler; the English-translation requirement is flagged as IMPORTANT where an agent will see it. The parenthetical '(Template 1)' and '(such as ...)' lists are slightly loose but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 11 parameters, 0% schema coverage, and no annotations, the description is materially incomplete. An output schema exists so return values need not be explained, but the vast majority of input parameters remain undefined, which is a significant gap for a code-generation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 11 parameters, so the description must compensate. It mentions generate_test_file's conditional effect, the English requirement, and lists a few field names (task, description, constraints, edge_cases), but leaves signature, language, test_cases, context, integration_note, and dependencies_allowed with no semantic explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate code for a standard function spec') and names the template variant, which helps distinguish it from generate_multi_function_module and generate_bugfix_refactor. It never explicitly contrasts itself with those siblings, so the differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the conditional effect of generate_test_file but gives no guidance on when to choose this tool over generate_multi_function_module or generate_bugfix_refactor. There are no prerequisites or exclusions, leaving the selection decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_project_testsA

Automatically detects and runs the project's test suite, returning the command's stdout and stderr. Can be used by the AI to verify correct behavior of generated code.

If 'custom_command' is provided, it runs that command line instead of auto-detecting.

ParametersJSON Schema
NameRequiredDescriptionDefault
custom_commandNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the auto-detection behavior, the custom_command override, and that stdout/stderr are returned, but it says nothing about execution environment, timeouts, truncation, or the fact that this executes arbitrary commands with potential side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, and every sentence adds distinct information (what it does, why to use it, how the parameter changes behavior). No waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema means return values need not be explained, and the single optional parameter is covered. But for an execution tool with zero annotation coverage, the absence of any safety, timeout, or environment context leaves a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains that custom_command replaces auto-detection with the given command line. It does not clarify expected format (shell string, args, quoting), so it falls short of fully documenting the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: it auto-detects and runs the project's test suite and returns stdout/stderr. It is clearly distinct from the generation-oriented siblings (generate_*), though it never explicitly names or contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Can be used by the AI to verify correct behavior of generated code' gives one implied usage context, which is a reasonable hint. However there is no when-not guidance, no mention of alternatives, and no prerequisites (e.g., project must be initialized, tests must exist).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedgenerate_bugfix_refactor
    • First observedgenerate_codebase_context
    • First observedgenerate_multi_function_module
    • First observedgenerate_standard_function
    • First observedrun_project_tests

TDQS

B3.4/5.0

Scored across 5 tools

Disambiguation4/5

The four generate_* tools map to distinct templates/scenarios (standard function, codebase-aware function, bugfix/refactor, multi-function module), so an agent can usually pick correctly. However, generate_standard_function and generate_codebase_context both produce a single function and differ only by whether codebase context is used, which could cause occasional misselection.

Naming Consistency5/5

All names use snake_case with a consistent verb_noun pattern (generate_* plus run_project_tests). The generate_* family is predictable and self-describing.

Tool Count5/5

Five tools is well-scoped for a code-generation server: four generation templates plus a test runner, each earning its place. No redundant or filler tools.

Completeness4/5

The surface covers code generation across common scenarios and test execution/verification, with test suites co-generated. Minor gaps exist around applying/writing generated code back to disk or editing existing files, but core workflows are covered.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers