ipython2-mcp
Enables agents to run Python 2 code in isolated IPython sessions, with tools for installing dependencies, managing sessions, executing code, and tab completion.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ipython2-mcpStart a Python 2 session and run print 'hello'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ipython2-mcp
This project allows an agent to run Python 2 IPython sessions via an MCP.
The MCP server is written in TypeScript. Each session is a Python 2 broker process owning one IPython kernel process, driven over JSON-RPC on stdio.
Tools
tool | params | what it does |
|
| pip-installs the kernel dependencies into a Python 2.7 interpreter |
|
| starts a named session on the given Python 2.7 interpreter |
|
| stops a session and frees the name |
| — | one line per live session |
|
| runs code, returns the transcript |
|
| tab completion |
Introspection, namespace listing, and history go through IPython itself (var?, var??,
%whos, %history) rather than dedicated tools.
Related MCP server: Sandbox Agent
Setup
Install and build the server:
npm install
npm run buildThat produces dist/ipython2-mcp.js — a single self-contained file. Register it with your MCP
client as node <repo>/dist/ipython2-mcp.js.
The repo's .mcp.json already does that with a path relative to the repo root, so an MCP client
started there picks the server up with no further configuration. Run npm run build first:
dist/ is not committed.
To distribute, ship dist/ipython2-mcp.js together with the python/ directory beside it — those
scripts are run by the user's Python 2 interpreter, so they cannot be bundled into the JS.
Then point python_path at any Python 2.7 interpreter — a venv's interpreter is the intended use.
It needs the kernel dependencies, which the agent can install for itself with
ipython_install_deps, or you can do by hand:
<py2>/python -m pip install -r python/requirements.txtNo kernelspec registration is needed; the kernel is launched with that interpreter directly.
SKILL.md is an agent-facing guide to working in the session; install it into .claude/skills/
if you want your agent to have it.
Distribution
dist/ipython2-mcp.js is committed, so users need nothing but node — no npm install, no
TypeScript toolchain. Two things ship: that file and the python/ directory beside it. Everything
else in the repo is development-only.
If you change src/, rebuild and commit the bundle in the same commit:
npm run build # tsc -> build/, esbuild -> dist/ipython2-mcp.js
npm run check:bundle # fails if the committed bundle is stale; wire into CIpython/ must remain a sibling of dist/ — the server resolves the broker relative to its own
location. See ARCHITECTURE.md for the plugin layout and .mcp.json snippet.
Tests
npm test # TypeScript
<py2>/python -m pip install -r python/requirements-dev.txt
<py2>/python -m pytest python/tests # broker, against a real kernel
PYTHON2=<py2>/python npm run test:e2e # smoke, needs a build firstSee ARCHITECTURE.md for the full design and the reasoning behind it.
Available Tools
6 toolsipython_completeA
Tab-complete a partial expression against the session's live namespace, e.g. df.gr or os.path.jo. The cursor is taken to be at the end of the prefix. Use this to discover attributes and names that actually exist before running code.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Session to complete in. | |
| prefix | Yes | Partial expression to complete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure. It adds useful detail about cursor handling ('cursor is taken to be at the end of the prefix') and the live namespace nature of the lookup, implying a read-only operation. However, it does not explicitly state that the tool is non-destructive, nor does it describe what happens on failure (e.g., empty completions) or whether results are returned as a list. These gaps prevent a higher score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The purpose is front-loaded, examples clarify usage immediately, and every clause contributes either to purpose, context, or a behavioral nuance. This is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite being a simple two-parameter tool, the description omits critical information about the return value. Since there is no output schema, the agent must infer what the tool returns to use it correctly. The description does not mention whether completions are returned as a list, how errors are signaled, or any pagination/limit behavior. This makes the definition incomplete for an agent that needs to parse the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description enriches the `prefix` parameter with examples and explains its interpretation, but it does not add information about the `name` parameter beyond the schema's 'Session to complete in.' The added value is modest, so a 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb ('Tab-complete') and resource ('partial expression against the session's live namespace'), followed by concrete examples (`df.gr`, `os.path.jo`). This clearly distinguishes it from sibling tools (run, start, etc.) and leaves no ambiguity about the tool's core function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use it: 'Use this to discover attributes and names that actually exist before running code.' This implies a pre-execution exploration context and implicitly contrasts with running code (via `ipython_run`). However, it does not explicitly name alternatives or conditions to avoid using it, so it falls short of a strong 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ipython_install_depsA
Install the kernel dependencies (IPython, ipykernel, jupyter_client) into a Python 2.7 interpreter using its own python -m pip, making it usable with ipython_start. Call this when ipython_start reports missing dependencies. Requires network access and writes to that interpreter's site-packages.
| Name | Required | Description | Default |
|---|---|---|---|
| timeout_ms | No | Give up after this long. Default 300000. | |
| python_path | Yes | Absolute path to the Python 2.7 interpreter to install into. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full burden. It discloses that the operation writes to the interpreter's site-packages and requires network access, and it specifies the mechanism (using that interpreter's own python -m pip). This makes the mutating nature and prerequisites clear, though it does not discuss idempotency or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each with a distinct job: state the action and target, give the trigger condition, and list side effects/prerequisites. No redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter install tool with no output schema, the description covers purpose, trigger, side effects, and a prerequisite (network). It could mention idempotency or failure output, but nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both parameters. The description adds context that python_path's own pip is used, clarifying that the installation targets that interpreter's environment. It does not add detail on timeout_ms, but the schema covers its meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Install'), the exact packages (IPython, ipykernel, jupyter_client), and the target environment (Python 2.7 interpreter). This differentiates it from siblings like ipython_start and ipython_status, which run or check rather than install.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to call it: when ipython_start reports missing dependencies. It also implies the alternative is to run ipython_start once deps are present, giving a clear trigger condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ipython_runA
Run code in a session and return the output transcript, exactly as a terminal would show it. A Python exception is a normal result: you get the traceback. Executions within a session are strictly serial, so a second call while one is running is rejected. On timeout the cell is interrupted and you get the partial output. Do not hold long-running work in a cell -- start it in a thread from your own code. This is a real IPython shell: use obj? and obj?? for signatures and source, %whos to list the namespace, and %history for past input.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | Code to execute. May span multiple lines and use IPython magics. | |
| name | Yes | Session to run in. | |
| timeout_ms | No | Interrupt the cell after this long. Default 30000, maximum 600000. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses traceback-on-exception, terminal-like output, serial execution rejection, timeout interruption with partial output, and IPython-specific behaviors like `obj?`, `%whos`, and `%history`.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded, and every sentence adds essential operational information without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a code-execution tool with no output schema and no annotations, the description fully covers return format, error behavior, concurrency, timeout semantics, and long-running work guidance. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds extra meaning by explaining timeout behavior (interruption and partial output) and reinforcing that code may use IPython magics across multiple lines.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Run code in a session and return the output transcript.' It clearly distinguishes itself from lifecycle siblings like ipython_start, ipython_stop, and ipython_status by focusing on execution and output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit constraints on usage: executions are serial, a second call while running is rejected, and long-running work should be threaded rather than kept in a cell. It does not explicitly name alternatives among siblings, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ipython_startA
Start a named Python 2 IPython session. python_path must be an absolute path to a Python 2.7 interpreter with this project's requirements installed -- pass a virtualenv's interpreter to work inside that virtualenv. Multiple named sessions can run at once and are fully independent.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for the kernel. Defaults to the MCP server's cwd. | |
| env | No | Extra environment variables, merged over the inherited environment. Use for PYTHONPATH or settings-module variables that must be set before imports resolve. | |
| name | Yes | Session name, 1-64 characters of [A-Za-z0-9_-]. Must not already be in use. | |
| python_path | Yes | Absolute path to the Python 2.7 interpreter to run this session with. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses useful traits: sessions are named, multiple can run concurrently, and they are fully independent. But it does not describe the session lifecycle, whether a background process is spawned, what happens on startup failure, or cleanup behavior, which are important for a process-starting tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences with no filler. The core action is front-loaded, and the critical python_path constraint and session-independence semantics are stated efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers important preconditions and session independence, but there is no output schema and no annotations, so the description must also explain invocation outcomes and side effects. It does not mention what the caller receives after starting, how a session is later referenced, or how this relates to ipython_install_deps and ipython_stop. This is adequate but has clear gaps for a process-management tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by clarifying that python_path must be absolute, must point to a Python 2.7 interpreter with project requirements installed, and that a virtualenv interpreter can be used. This gives the agent actionable selection guidance not present in the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Start a named Python 2 IPython session') and a resource. It implies a persistent session, and the mention of multiple independent named sessions adds useful scope. However, it does not explicitly distinguish itself from sibling ipython_run, which could be ambiguously similar in an agent's view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives important usage context: the python_path must be an absolute path to a Python 2.7 interpreter with project requirements installed, and virtualenv interpreters are supported. It does not, however, explicitly say when to prefer this over ipython_run, ipython_stop, or ipython_status, nor does it name alternatives or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ipython_statusA
List every session with its state (idle, busy, or dead), kernel pid, Python version, working directory, uptime, and execution count. A dead session keeps its name until ipython_stop clears it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the core behavior: it lists session details and includes the notable trait that dead sessions retain their name until ipython_stop clears them. This adds behavioral context beyond a generic 'list' and helps the agent understand the session lifecycle. It does not mention side effects, but for a read-only status tool, this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary action and output fields are front-loaded in the first sentence, and the second sentence adds a crucial nuance about dead sessions. Every word earns its place, and the structure is clean and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately specifies the return content by listing all fields (state, kernel pid, Python version, working directory, uptime, execution count) and the edge case for dead sessions. There are no parameters to document, and the behavior is fully described. An agent has enough to call the tool and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (empty schema). The description does not need to explain parameter semantics. The baseline for 0 params is 4, and there is no missing information here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'List' with a clear resource 'every session' and enumerates the exact fields (state, kernel pid, Python version, working directory, uptime, execution count). It also clarifies the behavior for dead sessions, distinguishing it from siblings like ipython_start or ipython_run. This leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (to inspect sessions) but does not explicitly state when to use this tool versus alternatives, such as 'use ipython_start to create a session' or 'use ipython_stop to clear dead sessions.' There is no exclusionary guidance, so it relies on the agent to infer context. This meets the 'implied usage' level but lacks explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ipython_stopA
Stop a session, kill its kernel, and free the name. Always works, including while a cell is executing -- this is the escape hatch for a wedged session. The namespace is lost.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Session to stop. | |
| timeout_ms | No | Grace period before the kernel is killed. Default 5000. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the safety burden. It discloses the destructive consequences—kernel killed, name freed, namespace lost—and gives a reliability guarantee. It does not detail response behavior or permissions, but the irreversible effect is explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler. The primary action is front-loaded, followed by the key escape-hatch behavior and the critical consequence, with every clause earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive stop tool with no annotations and no output schema, the description covers when to use it, what it does, and the key side effect. Minor gaps remain around timeout edge cases and postconditions, but nothing prevents correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents name and timeout_ms. The tool description adds no extra parameter-level meaning beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action chain: stop a session, kill its kernel, and free the name. This clearly distinguishes it from the sibling lifecycle tools (start/status/run/complete) and identifies a specific resource being acted on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit context for use: it is the escape hatch for a wedged session and works even while a cell is executing. It does not explicitly name alternatives or when not to use it, but the intended scenario is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
ipython_complete - First observed
ipython_install_deps - First observed
ipython_run - First observed
ipython_start - First observed
ipython_status - First observed
ipython_stop
TDQS
Scored across 6 tools
Each tool maps to a distinct lifecycle action: dependency install, session start, stop, status, code execution, and tab-completion. There is no overlap or ambiguity between tool purposes.
All tools share the ipython_ prefix and use clear snake_case verb-based suffixes: install_deps, start, stop, status, run, complete. The naming pattern is fully consistent and predictable.
Six tools is well-scoped for a session-management server. Every tool covers a necessary part of the workflow without redundancy or bloat.
The toolset covers the full lifecycle: dependency setup, session creation, inspection, execution, completion, and cleanup. Restarting a session is possible by combining stop and start, so there are no obvious dead ends.
Maintenance
Related MCP Connectors
MCP registry & directory: search, find & install 31k+ MCP servers & tools. Catalog and marketplace.
Create, deploy, and operate MCP servers directly from your GitHub repositories.
Personal assistant MCP server with search, execute, packages, jobs, secrets, and integrations.
Provision an OpenCode workbench and MCP stack on any Linux box, local or over SSH.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides persistent IPython shell sessions per conversation with DataFrame-centric architecture, enabling stateful data analysis, CLI tool execution, and integration of external MCP servers within the same workspace context.23Apache 2.0
- AlicenseAqualityCmaintenanceEnables code execution in isolated Docker containers with persistent IPython, Node.js, or R kernels, supporting file import/export and cross-session transfers via MCP tools.6MIT
- AlicenseNot gradedqualityDmaintenanceMCP server that provides a persistent IPython kernel for executing Python code with pre-loaded numpy, pandas, and matplotlib, supporting stateful computation and inline visualizations.2MIT
- AlicenseNot gradedqualityDmaintenanceProvides a persistent Python REPL session as a tool for executing code, managing files, installing packages, and initializing projects via the MCP protocol.1MIT