Skip to main content
Glama

ipynb-mcp-server

English | 简体中文

npm version license: MIT node: >=22 CI

An MCP server that lets any AI agent read, edit and run local Jupyter notebooks — safely, with zero setup.

Who this is for

  • Students. A course hands you a .ipynb full of empty cells to fill in. Let the agent read the task, write the cells and run them — and because every edit carries an anchor, an edit that races the cell you are typing in JupyterLab fails loudly instead of overwriting you.

  • Data people. Your analysis is the notebook, and its outputs hold the plots and dataframes you care about. "Change this one cell and re-run just it" is the whole point: resume touches one cell, not the forty before it — and the images come back as images.

  • Teachers. An assignment already lives in a notebook: let the agent read what each student's cells actually produce. It can edit a submission, but never silently — writes are anchored, atomic, and always preceded by a rolling backup.

  • Anyone pointing an agent at real files. The interesting part is not that it can edit your notebook; it is that it can show you it did not corrupt it.

Related MCP server: JupyterMCP

The three things that go wrong without it

What happens today

What this server does instead

You let an agent edit your notebook while it is also open in JupyterLab, and it overwrites a cell you changed — you find out when the file is already broken.

Every source edit carries a compare-and-swap anchor (expected_source_hash or expected_text). If the file changed since the agent last read it, the edit fails without writing — and the error already contains the current hash, so one retry succeeds. Writes are atomic and always preceded by a rolling backup.

"Run cell 87" re-runs the 40-minute training cell at the top, or the !wget that pulls 2 GB.

mode='resume' runs only the target cells in the live kernel. No kernel alive? mode='replay' silently rebuilds state from cell 0, then runs just the target — and replayed_cell_indexes tells you exactly what was re-executed. notebook_kernel(start) + resume runs one cell with zero replay.

To use an agent at all you must first stand up JupyterLab, copy a URL, manage a token and keep it running.

Nothing to start. stdio, one line of config, no port, no token. The server talks to a Jupyter kernel directly and dies with your client.

Install

Nothing to clone, nothing to build. The server ships as a prebuilt npm package — pick one of these three:

How

Command

When to use it

Run on demand — recommended

npx -y ipynb-mcp-server --root /path/to/your/notebooks

You only need it inside an MCP client's config. Nothing is installed permanently; npx fetches the published package into its cache the first time.

Install globally

npm install -g ipynb-mcp-server then ipynb-mcp-server --root /path/to/your/notebooks

You want the command on PATH, or want to pin a version (ipynb-mcp-server@<version>).

From source

git clone https://github.com/3021244161/ipynb-mcp && cd ipynb-mcp && pnpm install && pnpm build

Only if you are changing the code — see Development.

Requirements: Node ≥ 22 (which brings npm and npx). Python is needed only at the moment a cell actually runs, and the server finds it itself — see Interpreter selection. The installer never runs pip install and never compiles anything.

Add it to your client (60 seconds)

# 1. it is a plain stdio server — nothing to install, nothing to start
npx -y ipynb-mcp-server --root /path/to/your/notebooks
# 2. now put the one-line config below into your client and restart it

No pip install, no JupyterLab, no port, no token.

How it compares

Getting started

Edits that cannot silently corrupt your file

Long jobs

Maintenance

ipynb-mcp-server (this project)

one npx line, no service

CAS anchor + atomic write + rolling backup on every edit

resume / replay / one-cell resume, plus stale-cell analysis

active (2026-10)

datalayer/jupyter-mcp-server (~1.3k★)

needs a running Jupyter Server + SERVER_URL + TOKEN (or Docker)

—

—

active (company-maintained)

jupyter-ai-contrib/jupyter-server-mcp

Jupyter Server extension: installed into a running server

—

—

active

jbeno/cursor-notebook-mcp (~160★)

install from PyPI/npx, operates on the file

—

—

unmaintained since 2025-11

jjsantos01/jupyter-notebook-mcp (~130★)

bridges a running Jupyter over WebSocket

—

—

unmaintained since 2025-04

— means "not promised in that project's own documentation". Every cell above states only what each project's documentation and repository show (star counts and last-push dates read from the GitHub API on 2026-10-06); this table deliberately makes no claim about what the others do not do.

Extras: stale-cell analysis (which outputs are now invalid because their inputs changed), image outputs as native MCP image blocks, background execution with polling for long runs, per-notebook kernel lifecycle management.

Client configuration

// Claude Code / Cursor / VS Code (generic MCP stdio config)
{
  "mcpServers": {
    "ipynb": {
      "command": "npx",
      "args": ["-y", "ipynb-mcp-server"],
      "env": { "IPYNB_ROOT": "/absolute/path/to/your/notebooks" }
    }
  }
}

dsh users: install the companion dsh-ipynb-mcp bundle (see dsh-ipynb-mcp/).

The server fences all paths to IPYNB_ROOT (or --root, or the process working directory). The fence refuses to start when the root is your home directory or a filesystem root — pass an explicit --root.

Configuration

Precedence: CLI flags > IPYNB_* environment variables > defaults. Boolean flags accept --no- prefixes.

CLI

Env

Default

Description

--root <dir>

IPYNB_ROOT

cwd

Root directory fence

--allow-outside-root

IPYNB_ALLOW_OUTSIDE_ROOT

false

Allow paths outside the root

--read-only

IPYNB_READ_ONLY

false

Only notebook_read and kernel status allowed

--images <auto|never|always>

IPYNB_IMAGES

auto

Image block policy (auto: images only on full-output reads and runs)

--python <path>

IPYNB_PYTHON

auto

Explicit interpreter (failure is final)

--kernel-idle-seconds <n>

IPYNB_KERNEL_IDLE_SECONDS

3600

Idle kernel reclamation

--exec-timeout-seconds <n>

IPYNB_EXEC_TIMEOUT_SECONDS

300

Per-cell timeout

--background-threshold-seconds <n>

IPYNB_BACKGROUND_THRESHOLD_SECONDS

30

Runs whose estimated upper bound (timeout_seconds × target cells) exceeds 10× this go background

--backup-keep <n>

IPYNB_BACKUP_KEEP

10

Rolling backups per notebook

--artifact-dir <dir>

IPYNB_ARTIFACT_DIR

platform cache

Where image artifacts are written

--inline-text-chars <n>

IPYNB_INLINE_TEXT_CHARS

20000

Text output truncation threshold

--preview-lines <n>

IPYNB_PREVIEW_LINES

12

Source preview lines

--max-images-per-call <n>

IPYNB_MAX_IMAGES_PER_CALL

20

Image blocks per tool call

--max-image-bytes <n>

IPYNB_MAX_IMAGE_BYTES

20971520

Max bytes per image

--log-level <level>

IPYNB_LOG_LEVEL

info

stderr log level

Startup failures (bad values, root does not exist / is your home dir / artifact dir unwritable) exit with code 2.

With the defaults, a single-cell run is executed synchronously (its upper bound is exactly the 10× cut-off); two or more cells, or a raised --exec-timeout-seconds, return a background run_id you poll with notebook_run_status. The multiplier is DEVIATIONS.md D-015.

Interpreter selection

When a notebook needs a kernel, the interpreter is resolved by candidate chain: --python → the notebook's own metadata.kernelspec argv → .venv/venv next to the notebook → python3/python on PATH. Every failed candidate is recorded; only if all fail does the tool error (with a ready-to-run pip install ipykernel command — the server never installs anything itself). A .venv that disagrees with the kernelspec produces a kernelspec_mismatch warning; pass --python to pin one explicitly.

Outputs, images and large numbers

An include_outputs: 'full' read projects every stored output into one of the OutputItem shapes of SPEC.md §5.4 and returns image items as native MCP image blocks. Two value shapes are worth their own paragraph, because real notebooks contain both.

Images. A notebook may store an image as plain base64, as the data:image/png;base64,… URL people paste, wrapped over several lines, or as an array of lines (nbformat's multi-line form). A full-output read, and a run, return all of those as a valid image block decoded to the bytes in your file, with image_index and artifact_path naming that block and the artifact written for it (when no block is returned — a summary read, --images=never, or past max_images_per_call — both stay null, per SPEC §4.4). A value that cannot be decoded — including an empty one — never fails the call: the item stays kind: "image" with bytes: 0, artifact_path: null, image_index: null and a text_fallback saying why, and the call carries an image_materialize_failed warning. The stored value itself is left exactly as it was; the block is built from the decoded bytes, so a value your notebook holds but no client could decode is a degraded image, not a failed read.

Large integers in application/json. nbformat puts no type constraint on a json value, and an integer JavaScript cannot represent exactly (anything past ±2^53, e.g. 2**64) cannot survive a JSON number channel unchanged. Such a value is preserved byte-for-byte in the file — a read/write round trip no longer rounds it — and the response reports it rather than pretending: whenever that output is returned in full (a full-output read, or a run), the item carries a warnings array and the call-level warnings[] gains one entry — code output_truncated, which is already in SPEC §7's closed table — whose message contains the exact digits. value holds the nearest double, because that is what the JSON channel itself can carry and what any JSON client would parse; the digits you need are in the warning and in the file.

Known limitations

  • Execution is arbitrary code execution. Point the root at directories you would let the agent write to; the fence is a path boundary, not a sandbox. Only run notebooks you can read.

  • Stale analysis is Python-only. Non-Python kernels (R, Julia…) work for read/edit/run but skip stale analysis (method: "skipped"). It also cannot see through globals()/locals()/exec/eval/setattr, attribute assignments (obj.attr = 1) or import *. When a cell fails to parse, the whole analysis degrades to a conservative regex pass (all confidences drop to low; the regex pass additionally misses tuple unpacking, annotated assignments, indented assignments and with … as, and may flag identifiers inside strings/comments).

  • Interactive widgets are unsupported (application/vnd.jupyter.widget-view+json degrades to unsupported).

  • One response has a size budget (default 8 MiB), because the client dies above 10 MiB. A tool result travels as a single JSON-RPC line, and the MCP SDK's reader rejects a line over 10 MiB by closing the connection — you would then see -32000 Connection closed, and every later call in that session answers Not connected: the session is lost, not the response. So the server degrades before reaching that: the largest text values are shortened and marked, whole outputs are dropped if needed, and images are withheld once they no longer fit. Every removal arrives with an output_truncated warning, and withheld image bytes stay reachable through the artifact_path the payload already carries. The budget is deliberately below the cliff, so responses between 8 and 10 MiB are truncated rather than sent whole — content your client could have accepted arrives shortened and marked instead. Raise --max-response-bytes (or IPYNB_MAX_RESPONSE_BYTES) if your client's buffer is genuinely larger, and prefer include_outputs: 'summary' or cell_indexes for very large notebooks (DEVIATIONS.md D-065, D-067).

  • Memory is proportional to notebook size, and a notebook can be larger than the default heap. Reading and running hold the document in memory, and a real 37.5 MiB notebook (the kind an xgboost tuning session produces, with SHAP plots and dataframes in its outputs) peaks around 0.9 GiB across a full notebook_run. That is comfortable in Node's default heap, and it is what it is after a fix that removed a 16-fold parser defect — an earlier release needed 2.2 GiB for the same file and died with FATAL ERROR: Ineffective mark-compacts near heap limit, which the client saw only as -32000 Connection closed and which took every kernel on that server down with it (DEVIATIONS.md D-059). If you have notebooks well beyond this size, raise the limit for the server, e.g. NODE_OPTIONS=--max-old-space-size=4096 in the client's environment block; the cost is linear in the file, so 100 MiB notebooks want a few GiB.

  • Image-heavy single executions are still bounded by the transport. The sidecar speaks one NDJSON line per response, and a line is capped at 64 MiB; since an exec_cell response carries every output's base64, a single cell producing more than roughly 64 MiB of base64 image data (e.g. several near-max_image_bytes figures) fails with a protocol error rather than returning the images. Lower max_image_bytes, split the cell, or read the images back through notebook_read. Tracked as DEVIATIONS.md D-017. A protocol error tears the whole sidecar down, so every kernel it hosted (for every notebook in that interpreter) dies with it: the next notebook_run rebuilds silently via replay, but a long training cell that had already finished in memory will not be re-run.

  • A kernel that dies while no cell is running is noticed on the next request, not immediately. The sidecar polls the kernel process while it is executing a cell; between cells it only learns of an external kill (OOM killer, taskkill) when the next call arrives. notebook_run probes kernel liveness before reusing a session, so that case becomes a silent replay/rebuild rather than a failure — but the kernel's in-memory state is gone at that point.

  • An interpreter that imports ipykernel but cannot host a kernel is a hard failure, not a fallback. The candidate chain picks an interpreter by probing import ipykernel; if the kernel then fails to start (a broken pyzmq build is the common real-world case), the run fails with kernel_died and the error detail carries the sidecar's last stderr lines plus the OS exit status (e.g. code=3221226505 (0xC0000409) = STATUS_STACK_BUFFER_OVERRUN). The server does not silently retry with another interpreter (DEVIATIONS.md D-030).

  • mode='auto' can re-run the cells before your target. With no live kernel — the normal state, since read/edit never start one — a notebook_run that names specific cells resolves to replay: it silently executes every code cell before the target to rebuild the state those cells define, discarding that output. On a notebook whose first cells download a dataset or train for an hour, notebook_run(cell_selector='87') re-runs all of it. The response tells you afterwards (mode_used: "replay" plus replayed_cell_indexes). To run exactly one cell, start a kernel first — notebook_kernel(action='start'), then notebook_run(mode='resume') — which replays nothing; mode='resume' without a live kernel fails cleanly with kernel_not_available rather than guessing. SPEC §4.7 rule 1 mandates the silent replay, and the warning-code table is closed, so this is documented rather than changed (DEVIATIONS.md D-062).

  • A timed-out cell also ends its kernel (SPEC §4.7 rule 6), so in-memory state accumulated there is lost; the next run rebuilds through replay (D-025). The shutdown is asynchronous: the response returns first, and the kernel process may live on until the interrupted cell finishes on its own — seconds to minutes for a long computation. The management command reports no kernel immediately, and the process is gone by the time it ends; nothing is orphaned.

  • On Windows, interrupting a running cell usually does not work, so a timeout relies on that shutdown instead. Interrupting a kernel needs a console event that a stdio MCP server has no console to deliver; a time.sleep(30) cell ignores the interrupt and runs to completion, while the tool has already returned exec_timeout and closed the kernel. Observed on Windows; other platforms are not verified here. The timeout response no longer waits for a reply the running cell cannot send, so it arrives at timeout_seconds + about 10 s (interrupt grace plus teardown; measured 10.2 s for a 2 s timeout) rather than at the cell's full duration (D-033).

  • Every write is validated before it lands — for the cells the write rewrites. The bytes about to be written are re-parsed, and the cells this write changed are checked against the nbformat rules this implementation could break; a violation aborts the write with selfcheck_failed instead of producing a file Jupyter would refuse. This is deliberately not a full schema validation. Content that was already in your file and is merely carried forward is preserved and reported as a warning (file_changed_externally, with the rule and cell in the message), never used to block an edit or a run (D-032, D-037) — which means a file that already contained such content will still not satisfy nbformat.validate after a successful edit or run. Fix or clear that content yourself; this tool will not rewrite your history. scripts/e2e-smoke.mjs re-checks a real edit+run with Python's own nbformat.validate.

  • The kernel's connection file lives in the OS temp directory and is removed on every exit path this process controls. A hard kill (SIGKILL, power loss) can leave one behind there; it is never written into your notebook directory (D-023), and its name is unpredictable and its mode 0600 (D-034). It carries that kernel's HMAC key, so treat a leftover file as sensitive.

  • One run per notebook at a time. A second concurrent notebook_run on the same notebook fails with kernel_busy instead of interleaving cell executions — including when it arrives in the gap between the first run's cells. Different notebooks run in parallel.

  • Byte-level fidelity is logical, not literal. Serialization normalizes \uXXXX escapes and number spellings, so untouched regions of a heavily-escaped notebook may show file-level diffs. What is preserved is every value: 100.0 may come back as 100, 2.0 as 2, 1e-05 as 0.00001 — the same number, and the same is true of Python's own json.dumps output, which is where most of these spellings come from. What is not tolerated is a changed value: a number JavaScript cannot hold is written back as the literal that was there, with a warning naming the exact digits (see "Outputs, images and large numbers" above), and the same is true of literals that overflow (1e400) or underflow (1e-400) a double. Rolling backups (<name>.<timestamp>.ipynb.bak) cover the rest.

  • An unknown tool argument is an error, not a silent default. Sending cell_selector to notebook_read (whose argument is cell_indexes) fails with invalid_arguments instead of quietly reading the whole notebook.

  • A write racing another program's write is retried, then reported. On Windows the OS reports "another process holds this file" and "two renames collided" identically; the transient case is retried for about 0.75 s before notebook_locked is returned, so a momentary collision no longer looks like a locked file (D-035).

  • No auto-creation of notebooks, no format conversion, no collaboration features.

Differences from the obvious alternatives

  • vs jupyter nbconvert --execute: whole-notebook batch execution with no way to resume state; every run replays everything.

  • vs running code through the shell: no notebook state, no outputs written back to the .ipynb, no images, no CAS protection on edits.

  • vs attaching to a Jupyter Server: needs a running service and token management; this is zero-config stdio.

Development

pnpm install
pnpm typecheck && pnpm lint && pnpm test          # unit (no Python needed)
pnpm test:integration                              # real kernels (needs ipykernel)
pnpm smoke                                         # real stdio server driven by a real MCP client
pnpm check:package                                 # what `npm pack` would ship
pnpm build

pnpm lint is more than a linter: it also runs the zero-dependency checkers scripts/check-format.mjs (tabs, trailing whitespace), scripts/check-indent.mjs (block structure, via the TypeScript parser) and scripts/check-docs.mjs, which keeps the documentation invariants honest (one docs/DEVIATIONS.md with unique, gap-free entry ids, and docs/OPEN_QUESTIONS.md still a verbatim copy of SPEC §12). pnpm smoke starts the built server as a real stdio process, drives it with the SDK's own client, and checks the end-to-end behaviours the unit suite cannot reach — 26 checks today, including a run and a read of a notebook holding a data: URL image (the shape that used to fail the whole tools/call), a timeout, a background run, and Python's own nbformat.validate on the file it wrote. pnpm check:package asserts the shape of the shipped package (no compiled Python, no sources, no test files, no scratch scripts) and proves its own judgment with a resident mutation matrix. Before publishing, pnpm check:release packs the tarball, installs it into an empty directory and drives the INSTALLED binary over real stdio (six tools, a real kernel, a real execution, clean exit) — the one step a release can get wrong while every repository gate stays green.

The unit and integration suites share one dedicated venv in the system temp directory (never in the repository; override the location with IPYNB_TEST_VENV) and never touch your interpreters. The unit suite uses it too, because its analyzer cases start a real sidecar — which is also why both suites run their files serially. Set IPYNB_TEST_PYTHON to a base interpreter that already has ipykernel. A venv this suite did not create is never deleted, and one it cannot use is removed only when it created it.

Security: what the kernel inherits. The sidecar and the kernel it starts inherit this server's full environment (PATH, HOME, proxies, tokens — anything your MCP client passed in), and executed cells can read it. Kernel processes are also not sandboxed in any way: notebook_run executes whatever the notebook says, with your user's privileges. Point --root at directories you would let the agent write to, and run notebooks you are willing to execute.

Implementation follows the frozen spec in SPEC.md; every deviation is recorded in docs/DEVIATIONS.md.

License

MIT — see LICENSE.

Available Tools

6 tools
notebook_editB

Edit notebook cells. Every source change requires a compare-and-swap anchor (expected_source_hash or expected_text); a mismatch fails the whole request without writing.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYes1..32 edit ops (replace_lines, insert_lines, replace_source, insert_cell, delete_cell, move_cell, set_cell_type, clear_outputs)
pathYesNotebook path
dry_runNoCompute everything but do not write. Default false
create_backupNoDefault true
expected_content_hashNoOptional optimistic-lock hash

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one genuinely important trait: a hash mismatch aborts the entire request atomically without partial writes. However, it says nothing about dry_run/create_backup defaults, permission needs, or concurrency scope beyond the anchor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and then the critical constraint. No filler, though the anchor sentence would be stronger if it used the real parameter names.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch mutation tool with no annotations and no output schema, atomicity is covered well, but the description omits dry_run/backup semantics and return behavior, and it references non-existent parameters, leaving real gaps an agent must resolve from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is nominally 100% (baseline 3), but the description names 'expected_source_hash' and 'expected_text' as required anchors, neither of which exists in the schema, while the schema's actual lock parameter is 'expected_content_hash' described as merely optional. This mismatch is actively misleading about which parameter to pass and whether it is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Edit notebook cells') that cleanly separates it from siblings like notebook_read and notebook_run. It stops short of naming the alternative tools, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a precondition (a CAS anchor is required for source changes) but never says when to choose this tool over notebook_read/notebook_run or how to structure a batch edit. No when-not guidance or alternative routing is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notebook_kernelC

Inspect or manage the kernels held for notebooks: status, start, shutdown, restart.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoNotebook path; required for start/shutdown/restart, ignored for status
actionYesOne of 'status' | 'start' | 'shutdown' | 'restart'

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose that shutdown/restart are destructive to kernel state, whether elevated permissions are needed, or what the response looks like. Only the benign 'path ignored for status' detail is conveyed, and even that is already in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with no filler, and the resource is stated before the action list. It is appropriately brief, though for a multi-action tool the brevity edges toward under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool includes mutating actions (shutdown, restart) but has no annotations and no output schema to convey safety or return information. The description does not compensate, so an agent cannot tell how risky or reversible these operations are.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema (including path being required for start/shutdown/restart and ignored for status). The description's action list adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (kernels held for notebooks) and enumerates the operations (status, start, shutdown, restart), which is clearer than most. It does not explicitly contrast itself with siblings like notebook_run or notebook_run_status, leaving the kernel-vs-run distinction to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It lists possible actions but gives no guidance on when to choose this tool over notebook_run, notebook_run_status, or notebook_run_cancel. No prerequisites or conditions are stated; the agent must infer usage from the action names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notebook_readA

Read a Jupyter notebook: cell index, type, source and existing outputs. Set include_outputs='full' to get a cell's outputs including images.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesNotebook path (absolute, or relative to the server root)
cell_indexesNo0-based cell indexes to read; omit for all cells. This is an integer array — do not pass a range string.
include_sourceNoSource detail: 'none' | 'preview' (default) | 'full'
include_outputsNoOutput detail: 'none' | 'summary' (default) | 'full'
expected_content_hashNoOptional optimistic-lock hash

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the return shape and one behavioral nuance (full outputs include images), which is genuine value. However it never states that this is a non-mutating/safe operation, nor anything about cost, size limits, or behavior when the notebook is stale relative to expected_content_hash.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, front-loaded with what the tool reads before the optional flag tip. Every clause earns its place and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey the return shape, and it does list cell index, type, source and outputs. What is missing is edge behavior: whether omitted cell_indexes means all cells (only in the schema), ordering, and how truncation or errors surface.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, cell_indexes, include_source, include_outputs and expected_content_hash, making 3 the baseline. The description reinforces one parameter's effect (include_outputs='full' returns images), adding modest meaning but no syntax or default information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (Jupyter notebook) and enumerates what is returned: cell index, type, source and existing outputs. It is easily distinguished from notebook_edit and notebook_run by the read verb, but it never names or contrasts with those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage guidance is a conditional tip: set include_outputs='full' to get outputs including images. There is no statement of when to prefer this over notebook_edit or notebook_run, nor any exclusion or prerequisite (e.g. must the kernel be idle?). Usage is implied by the read verb rather than explained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notebook_runA

Execute notebook cells. mode='resume' runs only the target cells in the live kernel; 'replay' silently rebuilds state from cell 0 first; 'full' re-runs everything.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoExecution mode: 'auto' (default) | 'resume' | 'replay' | 'full'
pathYesNotebook path
cell_selectorNoWhich code cells to run: 'all' (default), '3', '0-4', or a comma list like '0-4,7,9'. This is a string selector — do not pass an array.
create_backupNoDefault true
write_outputsNoWrite fresh outputs back to the .ipynb. Default true
timeout_secondsNoPer-cell timeout in seconds (1..86400; default from server config)
clear_outputs_beforeNoClear target cells' outputs before running. Default true
expected_content_hashNoOptional optimistic-lock hash

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose one important side effect: 'replay' *silently* rebuilds state from cell 0 first, which tells the agent about hidden re-execution. It omits other key traits such as whether the call blocks or returns asynchronously (the existence of run_status/run_cancel siblings hints at async but is never stated), permission requirements, and the fact that outputs are written back to the .ipynb by default.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action, then a compact enumeration of the modes. Every clause carries information and nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter execution tool with no annotations and no output schema, the description covers the most important behavioral axis (kernel state per mode) but leaves the agent guessing about blocking vs async execution, timeout behavior, backup/write-back defaults, and the optimistic-lock hash. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema cannot: it explains the runtime semantics of resume/replay/full rather than just listing enum values. It leaves 'auto', cell_selector, create_backup, write_outputs, timeout_seconds, and expected_content_hash entirely to the schema, which limits it below a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Execute notebook cells') and immediately distinguishes three execution strategies by their kernel-state semantics. It does not explicitly name the sibling tools it differs from (notebook_edit, notebook_read), but the action is unambiguous enough that an agent can separate it from them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The mode semantics imply when each is appropriate ('resume' keeps live state, 'replay' rebuilds from cell 0, 'full' re-runs everything), which is useful implied guidance. However, it never states when to choose one over another, what 'auto' resolves to, or when to pair the call with notebook_run_status/notebook_run_cancel for long executions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notebook_run_cancelB

Cancel a background notebook run started by notebook_run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesRun id returned by notebook_run

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Cancel' implies mutation, but nothing is said about whether cancellation is graceful or forced, whether partial outputs are kept, whether the call is idempotent, or what happens for an already-completed run — all material for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action, with no redundant or filler content. Every word contributes to identifying the target operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema this is minimally adequate: purpose and the source of run_id are clear. However, with no annotations the description should have covered at least the effect of cancellation and its behavior on already-finished or unknown runs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents run_id as 'Run id returned by notebook_run'. The description restates this provenance but adds no format, constraint, or example beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (cancel) and resource (background notebook run), and ties the target to notebook_run as its originator, so an agent can tell it apart from notebook_run_status or notebook_run itself. It does not explicitly name a sibling it must not be confused with, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'started by notebook_run' implies the prerequisite that a run must already exist and that only background runs are cancellable. There is no explicit guidance on checking notebook_run_status first, nor any statement of when cancellation is inappropriate (e.g. already-finished runs), so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notebook_run_statusA

Poll a background notebook run started by notebook_run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesRun id returned by notebook_run

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. 'Poll' plus 'background' usefully signals a non-blocking, repeatable status check on an async operation, but nothing is said about terminal states, error behavior, rate limits, or whether the eventual result is included.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, zero waste, with the essential information (poll, background, origin) front-loaded. Nothing is padded or repeated from the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter polling tool with no annotations and no output schema, the description is minimally adequate. It does not tell the agent what a poll returns or when the run is finished, which is the main thing an agent needs to decide whether to call again.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter with 100% schema description coverage ('Run id returned by notebook_run'), so the schema already fully documents it. The description adds no syntax or format detail beyond that, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (poll) and resource (background notebook run) and explicitly ties the run to its originator, 'started by notebook_run'. That distinguishes it from the start (notebook_run), cancel (notebook_run_cancel), and content siblings, though it never names them as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies usage by stating the run must have been 'started by notebook_run', so the agent knows this is a follow-up call. However, there is no guidance on polling cadence, when to stop polling, or when to prefer notebook_run_cancel instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observednotebook_edit
    • First observednotebook_kernel
    • First observednotebook_read
    • First observednotebook_run
    • First observednotebook_run_cancel
    • First observednotebook_run_status

TDQS

A3.6/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a clearly distinct action on a notebook: read, edit, run, poll run status, cancel run, and manage kernels. The run_status/run_cancel pair are explicitly scoped to background runs started by notebook_run, so there is no real misselection risk.

Naming Consistency5/5

All six tools use a uniform snake_case notebook_* prefix with a predictable verb/noun suffix (read, edit, run, run_status, run_cancel, kernel). The pattern is consistent throughout with no style deviations.

Tool Count5/5

Six tools is well-scoped for a notebook server, covering the core read/edit/execute workflow plus async run lifecycle and kernel management. No tool feels redundant or padded.

Completeness4/5

Read, edit, run, and kernel lifecycle are covered, including async run polling and cancellation. Minor gaps remain, such as creating a brand-new notebook file or deleting cells/notebooks explicitly, but these are largely workable via edit.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers