Skip to main content
Glama

ipython2-mcp

This project allows an agent to run Python 2 IPython sessions via an MCP.

The MCP server is written in TypeScript. Each session is a Python 2 broker process owning one IPython kernel process, driven over JSON-RPC on stdio.

Tools

tool

params

what it does

ipython_install_deps

python_path, timeout_ms?

pip-installs the kernel dependencies into a Python 2.7 interpreter

ipython_start

name, python_path, cwd?, env?

starts a named session on the given Python 2.7 interpreter

ipython_stop

name, timeout_ms?

stops a session and frees the name

ipython_status

—

one line per live session

ipython_run

name, code, timeout_ms?

runs code, returns the transcript

ipython_complete

name, prefix

tab completion

Introspection, namespace listing, and history go through IPython itself (var?, var??, %whos, %history) rather than dedicated tools.

Related MCP server: Sandbox Agent

Setup

Install and build the server:

npm install
npm run build

That produces dist/ipython2-mcp.js — a single self-contained file. Register it with your MCP client as node <repo>/dist/ipython2-mcp.js.

The repo's .mcp.json already does that with a path relative to the repo root, so an MCP client started there picks the server up with no further configuration. Run npm run build first: dist/ is not committed.

To distribute, ship dist/ipython2-mcp.js together with the python/ directory beside it — those scripts are run by the user's Python 2 interpreter, so they cannot be bundled into the JS.

Then point python_path at any Python 2.7 interpreter — a venv's interpreter is the intended use. It needs the kernel dependencies, which the agent can install for itself with ipython_install_deps, or you can do by hand:

<py2>/python -m pip install -r python/requirements.txt

No kernelspec registration is needed; the kernel is launched with that interpreter directly.

SKILL.md is an agent-facing guide to working in the session; install it into .claude/skills/ if you want your agent to have it.

Distribution

dist/ipython2-mcp.js is committed, so users need nothing but node — no npm install, no TypeScript toolchain. Two things ship: that file and the python/ directory beside it. Everything else in the repo is development-only.

If you change src/, rebuild and commit the bundle in the same commit:

npm run build          # tsc -> build/, esbuild -> dist/ipython2-mcp.js
npm run check:bundle   # fails if the committed bundle is stale; wire into CI

python/ must remain a sibling of dist/ — the server resolves the broker relative to its own location. See ARCHITECTURE.md for the plugin layout and .mcp.json snippet.

Tests

npm test                                                   # TypeScript
<py2>/python -m pip install -r python/requirements-dev.txt
<py2>/python -m pytest python/tests                        # broker, against a real kernel
PYTHON2=<py2>/python npm run test:e2e                      # smoke, needs a build first

See ARCHITECTURE.md for the full design and the reasoning behind it.

Available Tools

6 tools
ipython_completeA

Tab-complete a partial expression against the session's live namespace, e.g. df.gr or os.path.jo. The cursor is taken to be at the end of the prefix. Use this to discover attributes and names that actually exist before running code.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSession to complete in.
prefixYesPartial expression to complete.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden for behavioral disclosure. It adds useful detail about cursor handling ('cursor is taken to be at the end of the prefix') and the live namespace nature of the lookup, implying a read-only operation. However, it does not explicitly state that the tool is non-destructive, nor does it describe what happens on failure (e.g., empty completions) or whether results are returned as a list. These gaps prevent a higher score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The purpose is front-loaded, examples clarify usage immediately, and every clause contributes either to purpose, context, or a behavioral nuance. This is concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite being a simple two-parameter tool, the description omits critical information about the return value. Since there is no output schema, the agent must infer what the tool returns to use it correctly. The description does not mention whether completions are returned as a list, how errors are signaled, or any pagination/limit behavior. This makes the definition incomplete for an agent that needs to parse the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description enriches the `prefix` parameter with examples and explains its interpretation, but it does not add information about the `name` parameter beyond the schema's 'Session to complete in.' The added value is modest, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a specific verb ('Tab-complete') and resource ('partial expression against the session's live namespace'), followed by concrete examples (`df.gr`, `os.path.jo`). This clearly distinguishes it from sibling tools (run, start, etc.) and leaves no ambiguity about the tool's core function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states when to use it: 'Use this to discover attributes and names that actually exist before running code.' This implies a pre-execution exploration context and implicitly contrasts with running code (via `ipython_run`). However, it does not explicitly name alternatives or conditions to avoid using it, so it falls short of a strong 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipython_install_depsA

Install the kernel dependencies (IPython, ipykernel, jupyter_client) into a Python 2.7 interpreter using its own python -m pip, making it usable with ipython_start. Call this when ipython_start reports missing dependencies. Requires network access and writes to that interpreter's site-packages.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeout_msNoGive up after this long. Default 300000.
python_pathYesAbsolute path to the Python 2.7 interpreter to install into.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full burden. It discloses that the operation writes to the interpreter's site-packages and requires network access, and it specifies the mechanism (using that interpreter's own python -m pip). This makes the mutating nature and prerequisites clear, though it does not discuss idempotency or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each with a distinct job: state the action and target, give the trigger condition, and list side effects/prerequisites. No redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter install tool with no output schema, the description covers purpose, trigger, side effects, and a prerequisite (network). It could mention idempotency or failure output, but nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters. The description adds context that python_path's own pip is used, clarifying that the installation targets that interpreter's environment. It does not add detail on timeout_ms, but the schema covers its meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Install'), the exact packages (IPython, ipykernel, jupyter_client), and the target environment (Python 2.7 interpreter). This differentiates it from siblings like ipython_start and ipython_status, which run or check rather than install.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to call it: when ipython_start reports missing dependencies. It also implies the alternative is to run ipython_start once deps are present, giving a clear trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipython_runA

Run code in a session and return the output transcript, exactly as a terminal would show it. A Python exception is a normal result: you get the traceback. Executions within a session are strictly serial, so a second call while one is running is rejected. On timeout the cell is interrupted and you get the partial output. Do not hold long-running work in a cell -- start it in a thread from your own code. This is a real IPython shell: use obj? and obj?? for signatures and source, %whos to list the namespace, and %history for past input.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesCode to execute. May span multiple lines and use IPython magics.
nameYesSession to run in.
timeout_msNoInterrupt the cell after this long. Default 30000, maximum 600000.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses traceback-on-exception, terminal-like output, serial execution rejection, timeout interruption with partial output, and IPython-specific behaviors like `obj?`, `%whos`, and `%history`.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded, and every sentence adds essential operational information without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a code-execution tool with no output schema and no annotations, the description fully covers return format, error behavior, concurrency, timeout semantics, and long-running work guidance. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds extra meaning by explaining timeout behavior (interruption and partial output) and reinforcing that code may use IPython magics across multiple lines.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Run code in a session and return the output transcript.' It clearly distinguishes itself from lifecycle siblings like ipython_start, ipython_stop, and ipython_status by focusing on execution and output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit constraints on usage: executions are serial, a second call while running is rejected, and long-running work should be threaded rather than kept in a cell. It does not explicitly name alternatives among siblings, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipython_startA

Start a named Python 2 IPython session. python_path must be an absolute path to a Python 2.7 interpreter with this project's requirements installed -- pass a virtualenv's interpreter to work inside that virtualenv. Multiple named sessions can run at once and are fully independent.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory for the kernel. Defaults to the MCP server's cwd.
envNoExtra environment variables, merged over the inherited environment. Use for PYTHONPATH or settings-module variables that must be set before imports resolve.
nameYesSession name, 1-64 characters of [A-Za-z0-9_-]. Must not already be in use.
python_pathYesAbsolute path to the Python 2.7 interpreter to run this session with.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses useful traits: sessions are named, multiple can run concurrently, and they are fully independent. But it does not describe the session lifecycle, whether a background process is spawned, what happens on startup failure, or cleanup behavior, which are important for a process-starting tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences with no filler. The core action is front-loaded, and the critical python_path constraint and session-independence semantics are stated efficiently. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers important preconditions and session independence, but there is no output schema and no annotations, so the description must also explain invocation outcomes and side effects. It does not mention what the caller receives after starting, how a session is later referenced, or how this relates to ipython_install_deps and ipython_stop. This is adequate but has clear gaps for a process-management tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by clarifying that python_path must be absolute, must point to a Python 2.7 interpreter with project requirements installed, and that a virtualenv interpreter can be used. This gives the agent actionable selection guidance not present in the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Start a named Python 2 IPython session') and a resource. It implies a persistent session, and the mention of multiple independent named sessions adds useful scope. However, it does not explicitly distinguish itself from sibling ipython_run, which could be ambiguously similar in an agent's view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives important usage context: the python_path must be an absolute path to a Python 2.7 interpreter with project requirements installed, and virtualenv interpreters are supported. It does not, however, explicitly say when to prefer this over ipython_run, ipython_stop, or ipython_status, nor does it name alternatives or exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipython_statusA

List every session with its state (idle, busy, or dead), kernel pid, Python version, working directory, uptime, and execution count. A dead session keeps its name until ipython_stop clears it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the core behavior: it lists session details and includes the notable trait that dead sessions retain their name until ipython_stop clears them. This adds behavioral context beyond a generic 'list' and helps the agent understand the session lifecycle. It does not mention side effects, but for a read-only status tool, this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The primary action and output fields are front-loaded in the first sentence, and the second sentence adds a crucial nuance about dead sessions. Every word earns its place, and the structure is clean and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description adequately specifies the return content by listing all fields (state, kernel pid, Python version, working directory, uptime, execution count) and the edge case for dead sessions. There are no parameters to document, and the behavior is fully described. An agent has enough to call the tool and interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema coverage is 100% (empty schema). The description does not need to explain parameter semantics. The baseline for 0 params is 4, and there is no missing information here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'List' with a clear resource 'every session' and enumerates the exact fields (state, kernel pid, Python version, working directory, uptime, execution count). It also clarifies the behavior for dead sessions, distinguishing it from siblings like ipython_start or ipython_run. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (to inspect sessions) but does not explicitly state when to use this tool versus alternatives, such as 'use ipython_start to create a session' or 'use ipython_stop to clear dead sessions.' There is no exclusionary guidance, so it relies on the agent to infer context. This meets the 'implied usage' level but lacks explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ipython_stopA

Stop a session, kill its kernel, and free the name. Always works, including while a cell is executing -- this is the escape hatch for a wedged session. The namespace is lost.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSession to stop.
timeout_msNoGrace period before the kernel is killed. Default 5000.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the safety burden. It discloses the destructive consequences—kernel killed, name freed, namespace lost—and gives a reliability guarantee. It does not detail response behavior or permissions, but the irreversible effect is explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler. The primary action is front-loaded, followed by the key escape-hatch behavior and the critical consequence, with every clause earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive stop tool with no annotations and no output schema, the description covers when to use it, what it does, and the key side effect. Minor gaps remain around timeout edge cases and postconditions, but nothing prevents correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents name and timeout_ms. The tool description adds no extra parameter-level meaning beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise action chain: stop a session, kill its kernel, and free the name. This clearly distinguishes it from the sibling lifecycle tools (start/status/run/complete) and identifies a specific resource being acted on.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit context for use: it is the escape hatch for a wedged session and works even while a cell is executing. It does not explicitly name alternatives or when not to use it, but the intended scenario is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedipython_complete
    • First observedipython_install_deps
    • First observedipython_run
    • First observedipython_start
    • First observedipython_status
    • First observedipython_stop

TDQS

A4.3/5.0

Scored across 6 tools

Disambiguation5/5

Each tool maps to a distinct lifecycle action: dependency install, session start, stop, status, code execution, and tab-completion. There is no overlap or ambiguity between tool purposes.

Naming Consistency5/5

All tools share the ipython_ prefix and use clear snake_case verb-based suffixes: install_deps, start, stop, status, run, complete. The naming pattern is fully consistent and predictable.

Tool Count5/5

Six tools is well-scoped for a session-management server. Every tool covers a necessary part of the workflow without redundancy or bloat.

Completeness5/5

The toolset covers the full lifecycle: dependency setup, session creation, inspection, execution, completion, and cleanup. Restarting a session is possible by combining stop and start, so there are no obvious dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides persistent IPython shell sessions per conversation with DataFrame-centric architecture, enabling stateful data analysis, CLI tool execution, and integration of external MCP servers within the same workspace context.
    23
    Apache 2.0
  • A
    license
    A
    quality
    C
    maintenance
    Enables code execution in isolated Docker containers with persistent IPython, Node.js, or R kernels, supporting file import/export and cross-session transfers via MCP tools.
    6
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a persistent Python REPL session as a tool for executing code, managing files, installing packages, and initializing projects via the MCP protocol.
    1
    MIT