Code Generator MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Code Generator MCP ServerGenerate a Python function to calculate factorial"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Code Generator MCP Template Server
An MCP (Model Context Protocol) server built using Python's FastMCP framework. It exposes 4 precise, structured code-generation tools backed by prompt templates to generate production-grade, parsed code using local or cloud-based OpenAI-compatible APIs (such as llama-server, qwen-coder, coder-expert, or OpenAI's API).
๐ Features
Exposes 4 key MCP tools that use structured templates to instruct the model to think step-by-step and produce clean, executable code:
generate_standard_function(Template 1): Generates a standalone function based on constraints, edge cases, test cases, and external integration notes.generate_codebase_context(Template 2): Generates a function that respects and integrates with existing codebase structures and dependencies.generate_bugfix_refactor(Template 3): Focuses on refactoring or repairing current buggy implementations based on problem descriptions and test expectations.generate_multi_function_module(Template 4): Generates a multi-function module, validating that there are no circular dependencies or undefined functions.
โ๏ธ Parser & Guardrails
Markdown Stripper: Automatic code parsing (
extract_code_from_response) strips any markdown code blocks (```python) generated by instruct models, guaranteeing only raw executable code is returned.Reasoning Fallback: Correctly handles DeepSeek-style reasoning models or
llama-serverconfigurations where all output is redirected into thereasoning_contentfield instead ofcontent.Max Tokens Guardrail: Enforces a
2048token limit per request to prevent local model reasoning loops and timeouts.
Related MCP server: mcp_server_for_claudes_toolbox
๐ ๏ธ Configuration
Configure the server using command-line arguments or environment variables:
Setting | CLI Argument | Environment Variable | Default Value | Description |
API URL |
|
|
| OpenAI-compatible endpoint |
Model |
|
|
| The model name to target |
API Key |
|
| (Empty) | API token (optional for local endpoints) |
For security reasons, do not pass--api-key via command-line arguments as it will be visible in plain text in the host process table. Use the environment variables instead.
๐ฆ Installation & Setup
You can install and configure the server either automatically using the installation script or manually via Python.
Option 1: Quick Installation Script (Recommended)
The project includes an install.sh script that automatically builds a standalone executable using PyInstaller, installs it to ~/.local/bin/code-generator-mcp, and configures your target AI coding agent.
Run the script and specify your target agent:
chmod +x install.sh ./install.sh <agent_type>Supported
<agent_type>values:claude-desktop(Claude Desktop global configuration)claude-code(Claude CLI global configuration at~/.claude.json)cursor(Cursor editor global config at~/.cursor/mcp.json)codex(Codex agent global config at~/.codex/config.toml)github-copilot(VS Code workspace-specific config at.vscode/mcp.json)windsurf(Windsurf IDE configuration at~/.codeium/windsurf/mcp_config.json)zed(Zed editor config at~/.config/zed/settings.json)agy(Antigravity settings config at~/.gemini/settings.json)
(Optional) Customize the endpoint and model in the agent's configuration file or environment variables after installation.
Option 2: Manual Setup via Python
Prerequisites
Python 3.10+
Dependencies installed in virtual environment:
python -m venv .venv source .venv/bin/activate pip install -r requirements.txt
Running Locally
To run the MCP server directly over standard input/output (stdio):
python src/code_generator_mcp/server.py --api-url http://localhost:8008/v1 --model coder-expertTesting during development
You can use mcp dev (from MCP CLI) to test the server interactively in a development UI:
mcp dev src/code_generator_mcp/server.py -- --api-url http://localhost:8008/v1 --model coder-expert๐ Integration Setup
To manually use this server with your favorite MCP client (like Claude Desktop or Cursor):
Claude Desktop Configuration
Open your Claude Desktop config file (usually located at ~/.config/Claude/claude_desktop_config.json on Linux/macOS or %APPDATA%\Claude\claude_desktop_config.json on Windows) and add the following entry:
{
"mcpServers": {
"code-generator-mcp": {
"command": "/path/to/project/.venv/bin/python",
"args": [
"/path/to/project/src/code_generator_mcp/server.py"
],
"env": {
"CODE_GEN_API_URL": "http://localhost:8008/v1",
"CODE_GEN_MODEL": "coder-expert"
}
}
}
}๐งช Testing
The codebase includes a fully-featured unit and integration test suite using pytest. Run tests with:
.venv/bin/pytestAvailable Tools
5 toolsgenerate_bugfix_refactorC
Generate fixed or refactored code (Template 3) using coder expert model.
If generate_test_file is True, it will co-generate a matching unit test suite.
IMPORTANT: All parameters (such as 'task', 'problem', 'expected_behavior', 'constraints', etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| problem | Yes | ||
| language | Yes | ||
| test_cases | Yes | ||
| constraints | No | ||
| current_code | Yes | ||
| expected_behavior | Yes | ||
| generate_test_file | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It usefully discloses one real behavioral constraint (all text parameters must be supplied in English, with a translation instruction) and the test co-generation behavior, but it says nothing about permissions, output format, limits, or what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in the first sentence, followed by two focused notes. The English/translation block is somewhat verbose but carries genuinely actionable information, so it earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter generation tool with zero schema descriptions, no annotations, and only partial parameter coverage, the description is thin. The output schema existing means return values need no explanation, but the input semantics and usage routing are largely missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all 8 parameters, so the description must compensate but does not. It names a few fields (task, problem, expected_behavior, constraints, generate_test_file) only to impose the English requirement, without explaining what language, current_code, or test_cases mean or how task differs from problem.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate fixed or refactored code') and references 'Template 3' plus the coder expert model, so the agent understands the operation. However, it does not distinguish itself from siblings like generate_standard_function or generate_multi_function_module, leaving the boundary inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus the sibling generators (standard function, multi-function module). The only conditional behavior mentioned is generate_test_file co-generation, which is a parameter effect rather than a when-to-use rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_codebase_contextC
Generate code for a function with codebase context (Template 2) using coder expert model.
If generate_test_file is True, it will co-generate a matching unit test suite.
IMPORTANT: All parameters (such as 'task', 'description', 'constraints', etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| language | Yes | ||
| signature | Yes | ||
| test_cases | Yes | ||
| constraints | No | ||
| description | Yes | ||
| existing_code | Yes | ||
| interacts_with | No | ||
| generate_test_file | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose that setting generate_test_file co-generates a unit test suite and that a 'coder expert model' is used, plus a notable English-only input constraint. It says nothing about how existing_code is used, permissions, latency, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core capability is front-loaded in the first sentence, with the conditional test-generation behavior immediately after. The English-only block is verbose but earns its place as a hard invocation constraint; only the unexplained 'Template 2' reference is dead weight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, 6-required generation tool with zero schema coverage and no annotations, the description is far too thin. An output schema exists so return values need not be explained, but the caller is given no way to know what signature, existing_code, interacts_with, or test_cases should contain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the schema contributes no meaning and the description must compensate. It only names 'task', 'description', and 'constraints' as examples in the English-language note and explains generate_test_file; signature, existing_code, interacts_with, test_cases, and language are left entirely undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate code for a function with codebase context'), so an agent knows this produces code from existing code context. However, it does not differentiate itself from the close sibling generate_standard_function, and 'Template 2' is opaque internal jargon that carries no meaning for the caller.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this over generate_standard_function, generate_multi_function_module, or generate_bugfix_refactor, despite these being obvious alternatives. The only conditional behavior mentioned is the generate_test_file flag, which is not selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_multi_function_moduleB
Generate code for a multi-function module (Template 4) using coder expert model.
If generate_test_file is True, it will co-generate a matching unit test suite.
IMPORTANT: All parameters (such as 'task', 'module_purpose', 'context', 'functions' list with its descriptions/constraints, etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| context | No | ||
| language | Yes | ||
| functions | Yes | ||
| test_cases | Yes | ||
| dependencies | No | ||
| shared_types | No | ||
| export_format | No | ||
| module_purpose | Yes | ||
| generate_test_file | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It usefully discloses the coder expert model, optional co-generated unit test suite, and mandatory English parameter translation, but omits permissions, side effects, determinism, and other operational traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then states the conditional test-generation behavior, then the critical language requirement. Every sentence adds information and there is no redundant restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter code-generation tool with no annotations, the description is incomplete. It does not explain key parameters, usage relative to siblings, or behavioral constraints beyond English translation and optional test generation, even though output-schema existence means return values need not be described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Top-level schema description coverage is 0% across 10 parameters, so the description must compensate. It only names or hints at task, module_purpose, context, functions, and generate_test_file, while language, test_cases, dependencies, shared_types, export_format, and most function-spec fields remain unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: generate code for a multi-function module (Template 4) using a coder expert model. The phrase 'multi-function module' distinguishes it from generate_standard_function, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no guidance on when to choose this over generate_standard_function, generate_codebase_context, or generate_bugfix_refactor. The only conditional behavior mentioned is if generate_test_file is True, which is execution behavior rather than tool selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_standard_functionC
Generate code for a standard function spec (Template 1) using coder expert model.
If generate_test_file is True, it will co-generate a matching unit test suite.
IMPORTANT: All parameters (such as 'task', 'description', 'constraints', 'edge_cases', etc.) MUST be provided in English. If the user's prompt or request is in Vietnamese or another language, the calling agent must automatically translate the text of these parameter values into English before invoking this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| context | No | ||
| language | Yes | ||
| signature | Yes | ||
| edge_cases | No | ||
| test_cases | Yes | ||
| constraints | No | ||
| description | Yes | ||
| integration_note | No | ||
| generate_test_file | No | ||
| dependencies_allowed | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that a coder expert model is used, that a test suite is co-generated when generate_test_file is true, and that all parameter values must be English. However it omits mutation/return behavior and error conditions, so the behavioral profile is only partially covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the English-translation requirement is flagged as IMPORTANT where an agent will see it. The parenthetical '(Template 1)' and '(such as ...)' lists are slightly loose but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 11 parameters, 0% schema coverage, and no annotations, the description is materially incomplete. An output schema exists so return values need not be explained, but the vast majority of input parameters remain undefined, which is a significant gap for a code-generation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 11 parameters, so the description must compensate. It mentions generate_test_file's conditional effect, the English requirement, and lists a few field names (task, description, constraints, edge_cases), but leaves signature, language, test_cases, context, integration_note, and dependencies_allowed with no semantic explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate code for a standard function spec') and names the template variant, which helps distinguish it from generate_multi_function_module and generate_bugfix_refactor. It never explicitly contrasts itself with those siblings, so the differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the conditional effect of generate_test_file but gives no guidance on when to choose this tool over generate_multi_function_module or generate_bugfix_refactor. There are no prerequisites or exclusions, leaving the selection decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_project_testsA
Automatically detects and runs the project's test suite, returning the command's stdout and stderr. Can be used by the AI to verify correct behavior of generated code.
If 'custom_command' is provided, it runs that command line instead of auto-detecting.
| Name | Required | Description | Default |
|---|---|---|---|
| custom_command | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the auto-detection behavior, the custom_command override, and that stdout/stderr are returned, but it says nothing about execution environment, timeouts, truncation, or the fact that this executes arbitrary commands with potential side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, and every sentence adds distinct information (what it does, why to use it, how the parameter changes behavior). No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema means return values need not be explained, and the single optional parameter is covered. But for an execution tool with zero annotation coverage, the absence of any safety, timeout, or environment context leaves a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it explains that custom_command replaces auto-detection with the given command line. It does not clarify expected format (shell string, args, quoting), so it falls short of fully documenting the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: it auto-detects and runs the project's test suite and returns stdout/stderr. It is clearly distinct from the generation-oriented siblings (generate_*), though it never explicitly names or contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Can be used by the AI to verify correct behavior of generated code' gives one implied usage context, which is a reasonable hint. However there is no when-not guidance, no mention of alternatives, and no prerequisites (e.g., project must be initialized, tests must exist).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
generate_bugfix_refactor - First observed
generate_codebase_context - First observed
generate_multi_function_module - First observed
generate_standard_function - First observed
run_project_tests
TDQS
Scored across 5 tools
The four generate_* tools map to distinct templates/scenarios (standard function, codebase-aware function, bugfix/refactor, multi-function module), so an agent can usually pick correctly. However, generate_standard_function and generate_codebase_context both produce a single function and differ only by whether codebase context is used, which could cause occasional misselection.
All names use snake_case with a consistent verb_noun pattern (generate_* plus run_project_tests). The generate_* family is predictable and self-describing.
Five tools is well-scoped for a code-generation server: four generation templates plus a test runner, each earning its place. No redundant or filler tools.
The surface covers code generation across common scenarios and test execution/verification, with test suites co-generated. Minor gaps exist around applying/writing generated code back to disk or editing existing files, but core workflows are covered.
Maintenance
Related MCP Connectors
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Nifty's MCP server โ exposes tasks, projects, messages, and files as tools for AI agents.
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAutomatic MCP Server & OpenAI Tools Bridge for apcore. Converts apcore module registries into MCP tool definitions and OpenAI-compatible function calling formats with zero boilerplate.320 npm1Apache 2.0
- FlicenseNot gradedqualityDmaintenanceExposes a set of CLI tools (test generation, documentation generation, linting, test running, code search) to AI assistants via MCP, allowing them to perform these tasks through natural language.3-
- AlicenseNot gradedqualityBmaintenanceExposes opencode's coding agent and shell as MCP tools, enabling MCP-only AI clients to execute shell commands, manage files, and run agent sessions with async job handling.4 npm1MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to access unified development tools including code generation, documentation synchronization, test case rendering, and architecture graph queries through a single MCP server.-