Skip to main content
Glama
bazinga012

MCP Code Executor

MCP コードエグゼキューター

鍛冶屋のバッジ

MCP Code Executor は、LLM が指定された Python 環境内で Python コードを実行できるようにする MCP サーバーです。これにより、LLM は環境で定義されたライブラリや依存関係にアクセスしながらコードを実行できます。また、トークン制限を超える可能性のある大規模なコードブロックを処理するための増分コード生成もサポートしています。

特徴

  • LLMプロンプトからPythonコードを実行する

  • トークン制限を克服するための増分コード生成のサポート

  • 指定された環境(Conda、virtualenv、または UV virtualenv)内でコードを実行する

  • 必要に応じて依存関係をインストールする

  • パッケージがすでにインストールされているかどうかを確認する

  • 実行時に環境を動的に構成する

  • 設定可能なコード保存ディレクトリ

Related MCP server: LLM Python Code Sandbox

前提条件

  • Node.jsがインストールされている

  • 次のいずれか:

    • Conda がインストールされ、必要な Conda 環境が作成されました

    • Python仮想環境

    • UV仮想環境

設定

  1. このリポジトリをクローンします:

git clone https://github.com/bazinga012/mcp_code_executor.git
  1. プロジェクト ディレクトリに移動します。

cd mcp_code_executor
  1. Node.js の依存関係をインストールします。

npm install
  1. プロジェクトをビルドします。

npm run build

構成

MCP Code Executor サーバーを構成するには、MCP サーバー構成ファイルに次の行を追加します。

Node.jsの使用

{
  "mcpServers": {
    "mcp-code-executor": {
      "command": "node",
      "args": [
        "/path/to/mcp_code_executor/build/index.js" 
      ],
      "env": {
        "CODE_STORAGE_DIR": "/path/to/code/storage",
        "ENV_TYPE": "conda",
        "CONDA_ENV_NAME": "your-conda-env"
      }
    }
  }
}

Dockerの使用

{
  "mcpServers": {
    "mcp-code-executor": {
      "command": "docker",
      "args": [
        "run",
        "-i",
        "--rm",
        "mcp-code-executor"
      ]
    }
  }
}

注: Dockerfileはvenv-uv環境タイプでのみテストされています。他の環境タイプでは追加の設定が必要になる場合があります。

環境変数

必須変数

  • CODE_STORAGE_DIR : 生成されたコードが保存されるディレクトリ

環境タイプ(1つの設定を選択)

  • Condaの場合:

    • ENV_TYPE : condaに設定

    • CONDA_ENV_NAME : 使用するConda環境の名前

  • 標準 Virtualenv の場合:

    • ENV_TYPE : venvに設定

    • VENV_PATH : 仮想環境ディレクトリへのパス

  • UV Virtualenvの場合:

    • ENV_TYPE : venv-uvに設定

    • UV_VENV_PATH : UV仮想環境ディレクトリへのパス

利用可能なツール

MCP コード エグゼキュータは、LLM に次のツールを提供します。

1. execute_code

設定された環境でPythonコードを実行します。短いコードスニペットに最適です。

{
  "name": "execute_code",
  "arguments": {
    "code": "import numpy as np\nprint(np.random.rand(3,3))",
    "filename": "matrix_gen"
  }
}

2. install_dependencies

環境に Python パッケージをインストールします。

{
  "name": "install_dependencies",
  "arguments": {
    "packages": ["numpy", "pandas", "matplotlib"]
  }
}

3. check_installed_packages

環境にパッケージがすでにインストールされているかどうかを確認します。

{
  "name": "check_installed_packages",
  "arguments": {
    "packages": ["numpy", "pandas", "non_existent_package"]
  }
}

4. configure_environment

環境構成を動的に変更します。

{
  "name": "configure_environment",
  "arguments": {
    "type": "conda",
    "conda_name": "new_env_name"
  }
}

5. get_environment_config

現在の環境構成を取得します。

{
  "name": "get_environment_config",
  "arguments": {}
}

6. initialize_code_file

初期コンテンツを含む新しいPythonファイルを作成します。トークン制限を超える可能性のある長いコードの最初のステップとしてこれを使用してください。

{
  "name": "initialize_code_file",
  "arguments": {
    "content": "def main():\n    print('Hello, world!')\n\nif __name__ == '__main__':\n    main()",
    "filename": "my_script"
  }
}

7. append_to_code_file

既存のPythonコードファイルにコンテンツを追加します。initialize_code_fileで作成されたファイルにコードを追加する場合に使用します。

{
  "name": "append_to_code_file",
  "arguments": {
    "file_path": "/path/to/code/storage/my_script_abc123.py",
    "content": "\ndef another_function():\n    print('This was appended to the file')\n"
  }
}

8. execute_code_file

既存のPythonファイルを実行します。initialize_code_fileとappend_to_code_fileでコードをビルドした後の最終ステップとして使用してください。

{
  "name": "execute_code_file",
  "arguments": {
    "file_path": "/path/to/code/storage/my_script_abc123.py"
  }
}

9. read_code_file

既存のPythonコードファイルの内容を読み取ります。ファイルの内容を追加したり実行したりする前に、ファイルの現在の状態を確認するために使用します。

{
  "name": "read_code_file",
  "arguments": {
    "file_path": "/path/to/code/storage/my_script_abc123.py"
  }
}

使用法

MCP コード エグゼキュータが設定されると、指定されたCODE_STORAGE_DIRにファイルを生成し、設定された環境内で実行することで、LLM が Python コードを実行できるようになります。

LLM はプロンプトでこの MCP サーバーを参照してコードを生成および実行できます。

大きなコードブロックの処理

LLM トークン制限を超える可能性のある大きなコード ブロックの場合は、増分コード生成アプローチを使用します。

  1. 基本構造を持つファイルをinitialize_code_fileを使用して初期化する

  2. append_to_code_fileを使用して後続の呼び出しにコードを追加します。

  3. 必要に応じてread_code_fileを使用してファイルの内容を確認します。

  4. execute_code_fileを使用して完全なコードを実行します。

このアプローチにより、LLM はトークンの制限に遭遇することなく、複雑な複数部分から成るコードを記述できます。

下位互換性

このパッケージは以前のバージョンとの下位互換性を維持しています。以前のバージョンでConda環境のみを指定していたユーザーは、設定を変更することなく引き続き作業できます。

貢献

貢献を歓迎します!問題を報告したり、プルリクエストを送信してください。

ライセンス

このプロジェクトは MIT ライセンスに基づいてライセンスされています。

Available Tools

9 tools
append_to_code_fileA

Append content to an existing Python code file. Use this to add more code to a file created with initialize_code_file, allowing you to build up larger code bases in parts.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesFull path to the file
contentYesContent to append to the file

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. While it mentions the tool appends content, it doesn't disclose important behavioral traits like whether it requires file existence, handles file permissions, what happens on errors, or if it overwrites existing content. For a mutation tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly concise with two sentences that each earn their place. The first sentence states the core purpose, and the second provides usage context and sibling tool reference. There's zero wasted text, and information is front-loaded appropriately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given this is a mutation tool with no annotations and no output schema, the description should do more to explain behavioral aspects. While it clearly states the purpose and usage guidelines, it lacks information about error conditions, file requirements, or what happens when the operation completes. For a tool that modifies files, this leaves important contextual gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (file_path and content) adequately. The description doesn't add any parameter-specific information beyond what's in the schema, such as file path format requirements or content encoding details. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Append content') and target resource ('existing Python code file'), distinguishing it from sibling tools like initialize_code_file (creates new files) and read_code_file (reads without modifying). It provides a precise verb+resource combination with clear scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('to add more code to a file created with initialize_code_file') and provides a clear alternative context ('allowing you to build up larger code bases in parts'). It directly references a sibling tool for comparison, giving clear guidance on appropriate usage scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_installed_packagesC

Check if packages are installed in the conda environment

ParametersJSON Schema
NameRequiredDescriptionDefault
packagesYesList of packages to check

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool checks installation status but doesn't describe what the output looks like (e.g., boolean per package, version details), error handling, or performance characteristics. This leaves significant gaps for an agent to understand how to interpret results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose with zero waste. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., success/failure indicators, detailed package info), which is critical for a check operation. For a tool with no structured output documentation, the description should compensate more.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the single parameter 'packages' clearly documented as 'List of packages to check'. The description adds no additional meaning beyond this, such as format examples (e.g., package names with versions) or constraints. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Check') and target ('packages are installed in the conda environment'), providing a specific verb+resource combination. However, it doesn't explicitly differentiate from sibling tools like 'get_environment_config' or 'install_dependencies', which might also provide package-related information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., whether a conda environment must be active), exclusions, or comparisons to sibling tools like 'get_environment_config' that might offer broader environment information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

configure_environmentC

Change the environment configuration settings

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYesType of Python environment
conda_nameNoName of the conda environment (required if type is 'conda')
venv_pathNoPath to the virtualenv (required if type is 'venv')
uv_venv_pathNoPath to the UV virtualenv (required if type is 'venv-uv')

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states 'Change' implying a mutation, but doesn't specify permissions needed, whether changes are reversible, potential side effects, or error handling. This is inadequate for a configuration tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of environment configuration, no annotations, and no output schema, the description is insufficient. It lacks details on behavioral traits, usage context, and expected outcomes, making it incomplete for effective tool invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all parameters. The description adds no additional meaning beyond the schema's details about environment types and paths, meeting the baseline for high coverage without extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Change the environment configuration settings' clearly states the action ('Change') and resource ('environment configuration settings'), making the purpose understandable. However, it doesn't differentiate from sibling tools like 'get_environment_config' (which likely reads rather than changes settings), leaving room for improvement in distinguishing functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, timing, or how it relates to sibling tools such as 'get_environment_config' or 'install_dependencies', leaving the agent without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_codeA

Execute Python code in the conda environment. For short code snippets only. For longer code, use initialize_code_file and append_to_code_file instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesPython code to execute
filenameNoOptional: Name of the file to save the code (default: generated UUID)

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions the environment ('conda environment') and a constraint on code length, but lacks details on execution behavior (e.g., timeout, output handling, error propagation) or safety considerations. It adds some context but is incomplete for a code execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero waste: the first states the purpose and constraint, the second provides alternative guidance. It is front-loaded with essential information and appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description covers basic purpose and usage but lacks details on execution behavior, return values, or error handling. It is minimally viable for a code execution tool but has clear gaps in contextual information needed for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters ('code' and 'filename'). The description does not add any meaning beyond the schema, such as explaining what 'short code snippets' entail or how the filename is used. Baseline 3 is appropriate as the schema handles parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Execute Python code') and resource ('in the conda environment'), and explicitly distinguishes it from sibling tools by mentioning 'initialize_code_file and append_to_code_file' for longer code, making the purpose unambiguous and differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool ('For short code snippets only') and when to use alternatives ('For longer code, use initialize_code_file and append_to_code_file instead'), offering clear context and exclusions without being misleading.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_code_fileA

Execute an existing Python file. Use this as the final step after building up code with initialize_code_file and append_to_code_file.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesFull path to the Python file to execute

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions that it executes a Python file, implying mutation/runtime effects, but lacks details on permissions, safety (e.g., sandboxing), error handling, or output behavior. It adds some context about being a 'final step' but misses key behavioral traits for an execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and followed by usage guidance. Every sentence earns its place with no wasted words, making it highly efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (executing code, which can have side effects), lack of annotations, and no output schema, the description is incomplete. It covers purpose and workflow but omits critical details like execution environment, return values, or error conditions, leaving gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the 'file_path' parameter fully. The description does not add any meaning beyond what the schema provides (e.g., format examples or constraints), resulting in a baseline score of 3 as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Execute') and resource ('an existing Python file'), distinguishing it from siblings like 'execute_code' (which likely executes code directly) and 'read_code_file' (which only reads). It directly addresses what the tool does without being vague or tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use this tool ('as the final step after building up code with initialize_code_file and append_to_code_file'), providing clear context and naming specific alternatives (siblings) for the workflow. This gives strong guidance on its role versus other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_environment_configB

Get the current environment configuration

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states 'Get' implies a read operation, but doesn't specify what 'environment configuration' includes (e.g., variables, paths, dependencies), whether it requires permissions, if it's cached or real-time, or what happens on errors. For a tool with zero annotation coverage, this leaves significant gaps in understanding its behavior beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence: 'Get the current environment configuration.' It's front-loaded with the core action and resource, with no wasted words or redundant information. This is appropriately sized and efficient for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (0 parameters, no output schema, no annotations), the description is minimally adequate. It states what the tool does but lacks details on output format, error handling, or how it differs from siblings. Without annotations or output schema, more context on return values or behavioral traits would improve completeness, but it's not entirely incomplete for such a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, meaning no parameters are documented in the schema. The description doesn't add parameter details, which is appropriate since there are none to explain. This meets the baseline of 4 for tools with zero parameters, as there's no need to compensate for missing schema information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get the current environment configuration' clearly states the verb ('Get') and resource ('environment configuration'), making the purpose understandable. However, it doesn't distinguish this tool from potential sibling tools like 'configure_environment' or 'check_installed_packages' that might also interact with environment settings, leaving some ambiguity about its specific scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With siblings like 'configure_environment' (which might modify settings) and 'check_installed_packages' (which might list installed components), there's no indication of whether this tool is for read-only access, current runtime settings, or other specific contexts. It lacks explicit when/when-not instructions or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

initialize_code_fileA

Create a new Python file with initial content. Use this as the first step for longer code that may exceed token limits. Follow with append_to_code_file for additional code.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesInitial content to write to the file
filenameNoOptional: Name of the file (default: generated UUID)

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states this creates a new file (implying a write operation) and mentions token limit considerations, but doesn't specify file system permissions, error handling, or what happens if the file already exists. It adds some context but lacks comprehensive behavioral details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (two sentences) with zero wasted words. The first sentence states the core purpose, and the second provides crucial usage guidance. Every sentence earns its place and is front-loaded with essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with no annotations and no output schema, the description does well by explaining the tool's role in a multi-step workflow and referencing its sibling. However, it doesn't mention what the tool returns (e.g., success confirmation, file path) or potential error conditions, leaving some gaps in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema (e.g., it doesn't explain content formatting or filename conventions). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Create a new Python file with initial content') and distinguishes it from its sibling 'append_to_code_file' by positioning it as 'the first step for longer code.' It explicitly names the resource (Python file) and verb (create), avoiding tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool ('as the first step for longer code that may exceed token limits') and when to use an alternative ('Follow with append_to_code_file for additional code'). It clearly differentiates usage contexts between initialization and appending.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_dependenciesC

Install Python dependencies in the conda environment

ParametersJSON Schema
NameRequiredDescriptionDefault
packagesYesList of packages to install

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action ('install') but doesn't reveal critical traits such as whether this requires admin permissions, if it's idempotent, potential side effects on the environment, or error handling. This leaves significant gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero wasted words. It's front-loaded with the core purpose, making it easy to parse quickly, which is ideal for conciseness in tool descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that performs installation (a mutation operation) with no annotations and no output schema, the description is inadequate. It lacks details on behavior, error cases, or what success looks like, leaving the agent under-informed about how to use it effectively in context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, with the 'packages' parameter fully documented in the schema. The description doesn't add any meaning beyond what the schema provides (e.g., package format examples or installation options), so it meets the baseline for high schema coverage without extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('install') and target ('Python dependencies in the conda environment'), providing a specific verb+resource combination. However, it doesn't explicitly distinguish this tool from sibling tools like 'check_installed_packages' or 'configure_environment', which prevents a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'check_installed_packages' for verification or 'configure_environment' for setup. There's no mention of prerequisites, typical use cases, or exclusions, leaving the agent with minimal contextual direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_code_fileA

Read the content of an existing Python code file. Use this to verify the current state of a file before appending more content or executing it.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesFull path to the file to read

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It correctly identifies this as a read operation and mentions the file must be 'existing,' but doesn't disclose error handling, file size limitations, encoding considerations, or what happens with non-existent files. The description provides basic behavioral context but lacks important operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two well-structured sentences that efficiently convey purpose and usage guidelines. Every word serves a clear function, with no redundant information or unnecessary elaboration. It's appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simple single-parameter read operation with no output schema, the description provides adequate context about when to use it and what it does. However, without annotations or output schema, it could benefit from more detail about return format, error conditions, or performance characteristics for a more complete picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'file_path' clearly documented as 'Full path to the file to read.' The description doesn't add any additional parameter semantics beyond what the schema provides, so it meets the baseline expectation when schema coverage is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Read the content') and target resource ('an existing Python code file'), distinguishing it from siblings like append_to_code_file or execute_code_file. It provides a precise verb+resource combination that leaves no ambiguity about the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('to verify the current state of a file before appending more content or executing it'), providing clear context for its application. It distinguishes this read operation from potential write or execute operations performed by sibling tools, offering practical guidance on tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv1.0.0
    • First observedappend_to_code_file
    • First observedcheck_installed_packages
    • First observedconfigure_environment
    • First observedexecute_code
    • First observedexecute_code_file
    • First observedget_environment_config
    • First observedinitialize_code_file
    • First observedinstall_dependencies
    • First observedread_code_file

TDQS

A3.8/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a clearly distinct purpose with no ambiguity. For example, initialize_code_file, append_to_code_file, and read_code_file handle different file operations, while execute_code and execute_code_file target different execution methods. The descriptions explicitly differentiate tools like execute_code (for short snippets) versus the file-based workflow.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern using snake_case, such as append_to_code_file, check_installed_packages, and configure_environment. There are no deviations in naming style or convention across the set, making them predictable and readable.

Tool Count5/5

With 9 tools, the count is well-scoped for a code execution server. Each tool earns its place by covering distinct aspects like file management, environment configuration, dependency handling, and code execution, without being overly sparse or bloated.

Completeness5/5

The tool set provides complete coverage for the code execution domain, including CRUD-like operations for files (initialize, append, read), environment management (configure, get config, install dependencies), and execution (code snippets, files). There are no obvious gaps, and the workflow from file creation to execution is fully supported.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers