ipynb-mcp
Summary: ipynb-mcp is an MCP server that lets any AI agent read, edit, and execute local Jupyter notebooks safely over stdio — no JupyterLab or service required.
Read notebooks (
notebook_read): cell index, type, source and existing outputs;include_outputs='full'returns outputs including images as native MCP image blocks.Edit notebooks (
notebook_edit): replace/insert lines, replace source, insert/delete/move cells, set cell type, clear outputs — up to 32 ops per call, with optionaldry_runand backups.Edit safely: every source change needs a compare-and-swap anchor (
expected_source_hash/expected_text); a mismatch fails without writing, and atomic writes are always preceded by a rolling backup.Run cells (
notebook_run):resumeruns only target cells in a live kernel,replaysilently rebuilds state from cell 0 first,fullre-runs everything;autopicks one.Target selection & control:
cell_selectorstrings like'0-4,7', per-celltimeout_seconds,write_outputs,clear_outputs_before.Long jobs go background: runs over the estimated-upper-bound threshold return a
run_id, polled withnotebook_run_statusand stopped withnotebook_run_cancel.Manage kernels (
notebook_kernel): per-notebookstatus,start,shutdown,restart; idle kernels are reclaimed automatically.Safety fence: all paths confined to
IPYNB_ROOT/--root(refuses home or filesystem root), optional--read-onlymode, and optimistic locking on reads/runs.Extras: stale-cell analysis (which outputs are invalidated by changed inputs), image artifact paths, output truncation warnings, and never re-running long upstream cells when resuming a live kernel.
Caveat: execution is arbitrary code execution with your user's privileges — point the root at directories you would let the agent write to.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ipynb-mcprun cell 3 in analysis.ipynb and show me the output"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ipynb-mcp-server
English | 简体中文
An MCP server that lets any AI agent read, edit and run local Jupyter notebooks — safely, with zero setup.
Who this is for
Students. A course hands you a
.ipynbfull of empty cells to fill in. Let the agent read the task, write the cells and run them — and because every edit carries an anchor, an edit that races the cell you are typing in JupyterLab fails loudly instead of overwriting you.Data people. Your analysis is the notebook, and its outputs hold the plots and dataframes you care about. "Change this one cell and re-run just it" is the whole point:
resumetouches one cell, not the forty before it — and the images come back as images.Teachers. An assignment already lives in a notebook: let the agent read what each student's cells actually produce. It can edit a submission, but never silently — writes are anchored, atomic, and always preceded by a rolling backup.
Anyone pointing an agent at real files. The interesting part is not that it can edit your notebook; it is that it can show you it did not corrupt it.
Related MCP server: JupyterMCP
The three things that go wrong without it
What happens today | What this server does instead |
You let an agent edit your notebook while it is also open in JupyterLab, and it overwrites a cell you changed — you find out when the file is already broken. | Every source edit carries a compare-and-swap anchor ( |
"Run cell 87" re-runs the 40-minute training cell at the top, or the |
|
To use an agent at all you must first stand up JupyterLab, copy a URL, manage a token and keep it running. | Nothing to start. stdio, one line of config, no port, no token. The server talks to a Jupyter kernel directly and dies with your client. |
Install
Nothing to clone, nothing to build. The server ships as a prebuilt npm package — pick one of these three:
How | Command | When to use it |
Run on demand — recommended |
| You only need it inside an MCP client's config. Nothing is installed permanently; |
Install globally |
| You want the command on |
From source |
| Only if you are changing the code — see Development. |
Requirements: Node ≥ 22 (which brings npm and npx). Python is needed only at the moment a cell actually runs, and the server finds it itself — see Interpreter selection. The installer never runs pip install and never compiles anything.
Add it to your client (60 seconds)
# 1. it is a plain stdio server — nothing to install, nothing to start
npx -y ipynb-mcp-server --root /path/to/your/notebooks
# 2. now put the one-line config below into your client and restart itNo pip install, no JupyterLab, no port, no token.
How it compares
Getting started | Edits that cannot silently corrupt your file | Long jobs | Maintenance | |
ipynb-mcp-server (this project) | one | CAS anchor + atomic write + rolling backup on every edit |
| active (2026-10) |
datalayer/jupyter-mcp-server (~1.3k★) | needs a running Jupyter Server + | — | — | active (company-maintained) |
Jupyter Server extension: installed into a running server | — | — | active | |
jbeno/cursor-notebook-mcp (~160★) | install from PyPI/npx, operates on the file | — | — | unmaintained since 2025-11 |
jjsantos01/jupyter-notebook-mcp (~130★) | bridges a running Jupyter over WebSocket | — | — | unmaintained since 2025-04 |
—means "not promised in that project's own documentation". Every cell above states only what each project's documentation and repository show (star counts and last-push dates read from the GitHub API on 2026-10-06); this table deliberately makes no claim about what the others do not do.
Extras: stale-cell analysis (which outputs are now invalid because their inputs changed), image outputs as native MCP image blocks, background execution with polling for long runs, per-notebook kernel lifecycle management.
Client configuration
// Claude Code / Cursor / VS Code (generic MCP stdio config)
{
"mcpServers": {
"ipynb": {
"command": "npx",
"args": ["-y", "ipynb-mcp-server"],
"env": { "IPYNB_ROOT": "/absolute/path/to/your/notebooks" }
}
}
}dsh users: install the companion dsh-ipynb-mcp bundle (see dsh-ipynb-mcp/).
The server fences all paths to IPYNB_ROOT (or --root, or the process working directory). The fence refuses to start when the root is your home directory or a filesystem root — pass an explicit --root.
Configuration
Precedence: CLI flags > IPYNB_* environment variables > defaults. Boolean flags accept --no- prefixes.
CLI | Env | Default | Description |
|
| cwd | Root directory fence |
|
|
| Allow paths outside the root |
|
|
| Only |
|
|
| Image block policy ( |
|
| auto | Explicit interpreter (failure is final) |
|
|
| Idle kernel reclamation |
|
|
| Per-cell timeout |
|
|
| Runs whose estimated upper bound ( |
|
|
| Rolling backups per notebook |
|
| platform cache | Where image artifacts are written |
|
|
| Text output truncation threshold |
|
|
| Source preview lines |
|
|
| Image blocks per tool call |
|
|
| Max bytes per image |
|
|
| stderr log level |
Startup failures (bad values, root does not exist / is your home dir / artifact dir unwritable) exit with code 2.
With the defaults, a single-cell run is executed synchronously (its upper bound is exactly the 10× cut-off); two or more cells, or a raised --exec-timeout-seconds, return a background run_id you poll with notebook_run_status. The multiplier is DEVIATIONS.md D-015.
Interpreter selection
When a notebook needs a kernel, the interpreter is resolved by candidate chain: --python → the notebook's own metadata.kernelspec argv → .venv/venv next to the notebook → python3/python on PATH. Every failed candidate is recorded; only if all fail does the tool error (with a ready-to-run pip install ipykernel command — the server never installs anything itself). A .venv that disagrees with the kernelspec produces a kernelspec_mismatch warning; pass --python to pin one explicitly.
Outputs, images and large numbers
An include_outputs: 'full' read projects every stored output into one of the OutputItem shapes of SPEC.md §5.4 and returns image items as native MCP image blocks. Two value shapes are worth their own paragraph, because real notebooks contain both.
Images. A notebook may store an image as plain base64, as the data:image/png;base64,… URL people paste, wrapped over several lines, or as an array of lines (nbformat's multi-line form). A full-output read, and a run, return all of those as a valid image block decoded to the bytes in your file, with image_index and artifact_path naming that block and the artifact written for it (when no block is returned — a summary read, --images=never, or past max_images_per_call — both stay null, per SPEC §4.4). A value that cannot be decoded — including an empty one — never fails the call: the item stays kind: "image" with bytes: 0, artifact_path: null, image_index: null and a text_fallback saying why, and the call carries an image_materialize_failed warning. The stored value itself is left exactly as it was; the block is built from the decoded bytes, so a value your notebook holds but no client could decode is a degraded image, not a failed read.
Large integers in application/json. nbformat puts no type constraint on a json value, and an integer JavaScript cannot represent exactly (anything past ±2^53, e.g. 2**64) cannot survive a JSON number channel unchanged. Such a value is preserved byte-for-byte in the file — a read/write round trip no longer rounds it — and the response reports it rather than pretending: whenever that output is returned in full (a full-output read, or a run), the item carries a warnings array and the call-level warnings[] gains one entry — code output_truncated, which is already in SPEC §7's closed table — whose message contains the exact digits. value holds the nearest double, because that is what the JSON channel itself can carry and what any JSON client would parse; the digits you need are in the warning and in the file.
Known limitations
Execution is arbitrary code execution. Point the root at directories you would let the agent write to; the fence is a path boundary, not a sandbox. Only run notebooks you can read.
Stale analysis is Python-only. Non-Python kernels (R, Julia…) work for read/edit/run but skip stale analysis (
method: "skipped"). It also cannot see throughglobals()/locals()/exec/eval/setattr, attribute assignments (obj.attr = 1) orimport *. When a cell fails to parse, the whole analysis degrades to a conservative regex pass (all confidences drop tolow; the regex pass additionally misses tuple unpacking, annotated assignments, indented assignments andwith … as, and may flag identifiers inside strings/comments).Interactive widgets are unsupported (
application/vnd.jupyter.widget-view+jsondegrades tounsupported).One response has a size budget (default 8 MiB), because the client dies above 10 MiB. A tool result travels as a single JSON-RPC line, and the MCP SDK's reader rejects a line over 10 MiB by closing the connection — you would then see
-32000 Connection closed, and every later call in that session answersNot connected: the session is lost, not the response. So the server degrades before reaching that: the largest text values are shortened and marked, whole outputs are dropped if needed, and images are withheld once they no longer fit. Every removal arrives with anoutput_truncatedwarning, and withheld image bytes stay reachable through theartifact_paththe payload already carries. The budget is deliberately below the cliff, so responses between 8 and 10 MiB are truncated rather than sent whole — content your client could have accepted arrives shortened and marked instead. Raise--max-response-bytes(orIPYNB_MAX_RESPONSE_BYTES) if your client's buffer is genuinely larger, and preferinclude_outputs: 'summary'orcell_indexesfor very large notebooks (DEVIATIONS.mdD-065, D-067).Memory is proportional to notebook size, and a notebook can be larger than the default heap. Reading and running hold the document in memory, and a real 37.5 MiB notebook (the kind an xgboost tuning session produces, with SHAP plots and dataframes in its outputs) peaks around 0.9 GiB across a full
notebook_run. That is comfortable in Node's default heap, and it is what it is after a fix that removed a 16-fold parser defect — an earlier release needed 2.2 GiB for the same file and died withFATAL ERROR: Ineffective mark-compacts near heap limit, which the client saw only as-32000 Connection closedand which took every kernel on that server down with it (DEVIATIONS.mdD-059). If you have notebooks well beyond this size, raise the limit for the server, e.g.NODE_OPTIONS=--max-old-space-size=4096in the client's environment block; the cost is linear in the file, so 100 MiB notebooks want a few GiB.Image-heavy single executions are still bounded by the transport. The sidecar speaks one NDJSON line per response, and a line is capped at 64 MiB; since an
exec_cellresponse carries every output's base64, a single cell producing more than roughly 64 MiB of base64 image data (e.g. several near-max_image_bytesfigures) fails with a protocol error rather than returning the images. Lowermax_image_bytes, split the cell, or read the images back throughnotebook_read. Tracked asDEVIATIONS.mdD-017. A protocol error tears the whole sidecar down, so every kernel it hosted (for every notebook in that interpreter) dies with it: the nextnotebook_runrebuilds silently viareplay, but a long training cell that had already finished in memory will not be re-run.A kernel that dies while no cell is running is noticed on the next request, not immediately. The sidecar polls the kernel process while it is executing a cell; between cells it only learns of an external kill (OOM killer,
taskkill) when the next call arrives.notebook_runprobes kernel liveness before reusing a session, so that case becomes a silentreplay/rebuild rather than a failure — but the kernel's in-memory state is gone at that point.An interpreter that imports
ipykernelbut cannot host a kernel is a hard failure, not a fallback. The candidate chain picks an interpreter by probingimport ipykernel; if the kernel then fails to start (a broken pyzmq build is the common real-world case), the run fails withkernel_diedand the error detail carries the sidecar's last stderr lines plus the OS exit status (e.g.code=3221226505 (0xC0000409) = STATUS_STACK_BUFFER_OVERRUN). The server does not silently retry with another interpreter (DEVIATIONS.mdD-030).mode='auto'can re-run the cells before your target. With no live kernel — the normal state, since read/edit never start one — anotebook_runthat names specific cells resolves toreplay: it silently executes every code cell before the target to rebuild the state those cells define, discarding that output. On a notebook whose first cells download a dataset or train for an hour,notebook_run(cell_selector='87')re-runs all of it. The response tells you afterwards (mode_used: "replay"plusreplayed_cell_indexes). To run exactly one cell, start a kernel first —notebook_kernel(action='start'), thennotebook_run(mode='resume')— which replays nothing;mode='resume'without a live kernel fails cleanly withkernel_not_availablerather than guessing. SPEC §4.7 rule 1 mandates the silent replay, and the warning-code table is closed, so this is documented rather than changed (DEVIATIONS.mdD-062).A timed-out cell also ends its kernel (SPEC §4.7 rule 6), so in-memory state accumulated there is lost; the next run rebuilds through
replay(D-025). The shutdown is asynchronous: the response returns first, and the kernel process may live on until the interrupted cell finishes on its own — seconds to minutes for a long computation. The management command reports no kernel immediately, and the process is gone by the time it ends; nothing is orphaned.On Windows, interrupting a running cell usually does not work, so a timeout relies on that shutdown instead. Interrupting a kernel needs a console event that a stdio MCP server has no console to deliver; a
time.sleep(30)cell ignores the interrupt and runs to completion, while the tool has already returnedexec_timeoutand closed the kernel. Observed on Windows; other platforms are not verified here. The timeout response no longer waits for a reply the running cell cannot send, so it arrives attimeout_seconds+ about 10 s (interrupt grace plus teardown; measured 10.2 s for a 2 s timeout) rather than at the cell's full duration (D-033).Every write is validated before it lands — for the cells the write rewrites. The bytes about to be written are re-parsed, and the cells this write changed are checked against the nbformat rules this implementation could break; a violation aborts the write with
selfcheck_failedinstead of producing a file Jupyter would refuse. This is deliberately not a full schema validation. Content that was already in your file and is merely carried forward is preserved and reported as a warning (file_changed_externally, with the rule and cell in the message), never used to block an edit or a run (D-032, D-037) — which means a file that already contained such content will still not satisfynbformat.validateafter a successful edit or run. Fix or clear that content yourself; this tool will not rewrite your history.scripts/e2e-smoke.mjsre-checks a real edit+run with Python's ownnbformat.validate.The kernel's connection file lives in the OS temp directory and is removed on every exit path this process controls. A hard kill (SIGKILL, power loss) can leave one behind there; it is never written into your notebook directory (D-023), and its name is unpredictable and its mode 0600 (D-034). It carries that kernel's HMAC key, so treat a leftover file as sensitive.
One run per notebook at a time. A second concurrent
notebook_runon the same notebook fails withkernel_busyinstead of interleaving cell executions — including when it arrives in the gap between the first run's cells. Different notebooks run in parallel.Byte-level fidelity is logical, not literal. Serialization normalizes
\uXXXXescapes and number spellings, so untouched regions of a heavily-escaped notebook may show file-level diffs. What is preserved is every value:100.0may come back as100,2.0as2,1e-05as0.00001— the same number, and the same is true of Python's ownjson.dumpsoutput, which is where most of these spellings come from. What is not tolerated is a changed value: a number JavaScript cannot hold is written back as the literal that was there, with a warning naming the exact digits (see "Outputs, images and large numbers" above), and the same is true of literals that overflow (1e400) or underflow (1e-400) a double. Rolling backups (<name>.<timestamp>.ipynb.bak) cover the rest.An unknown tool argument is an error, not a silent default. Sending
cell_selectortonotebook_read(whose argument iscell_indexes) fails withinvalid_argumentsinstead of quietly reading the whole notebook.A write racing another program's write is retried, then reported. On Windows the OS reports "another process holds this file" and "two renames collided" identically; the transient case is retried for about 0.75 s before
notebook_lockedis returned, so a momentary collision no longer looks like a locked file (D-035).No auto-creation of notebooks, no format conversion, no collaboration features.
Differences from the obvious alternatives
vs
jupyter nbconvert --execute: whole-notebook batch execution with no way to resume state; every run replays everything.vs running code through the shell: no notebook state, no outputs written back to the
.ipynb, no images, no CAS protection on edits.vs attaching to a Jupyter Server: needs a running service and token management; this is zero-config stdio.
Development
pnpm install
pnpm typecheck && pnpm lint && pnpm test # unit (no Python needed)
pnpm test:integration # real kernels (needs ipykernel)
pnpm smoke # real stdio server driven by a real MCP client
pnpm check:package # what `npm pack` would ship
pnpm buildpnpm lint is more than a linter: it also runs the zero-dependency checkers scripts/check-format.mjs (tabs, trailing whitespace), scripts/check-indent.mjs (block structure, via the TypeScript parser) and scripts/check-docs.mjs, which keeps the documentation invariants honest (one docs/DEVIATIONS.md with unique, gap-free entry ids, and docs/OPEN_QUESTIONS.md still a verbatim copy of SPEC §12). pnpm smoke starts the built server as a real stdio process, drives it with the SDK's own client, and checks the end-to-end behaviours the unit suite cannot reach — 26 checks today, including a run and a read of a notebook holding a data: URL image (the shape that used to fail the whole tools/call), a timeout, a background run, and Python's own nbformat.validate on the file it wrote. pnpm check:package asserts the shape of the shipped package (no compiled Python, no sources, no test files, no scratch scripts) and proves its own judgment with a resident mutation matrix. Before publishing, pnpm check:release packs the tarball, installs it into an empty directory and drives the INSTALLED binary over real stdio (six tools, a real kernel, a real execution, clean exit) — the one step a release can get wrong while every repository gate stays green.
The unit and integration suites share one dedicated venv in the system temp directory (never in the repository; override the location with IPYNB_TEST_VENV) and never touch your interpreters. The unit suite uses it too, because its analyzer cases start a real sidecar — which is also why both suites run their files serially. Set IPYNB_TEST_PYTHON to a base interpreter that already has ipykernel. A venv this suite did not create is never deleted, and one it cannot use is removed only when it created it.
Security: what the kernel inherits. The sidecar and the kernel it starts inherit this server's full environment (PATH, HOME, proxies, tokens — anything your MCP client passed in), and executed cells can read it. Kernel processes are also not sandboxed in any way: notebook_run executes whatever the notebook says, with your user's privileges. Point --root at directories you would let the agent write to, and run notebooks you are willing to execute.
Implementation follows the frozen spec in SPEC.md; every deviation is recorded in docs/DEVIATIONS.md.
License
MIT — see LICENSE.
Available Tools
6 toolsnotebook_editB
Edit notebook cells. Every source change requires a compare-and-swap anchor (expected_source_hash or expected_text); a mismatch fails the whole request without writing.
| Name | Required | Description | Default |
|---|---|---|---|
| ops | Yes | 1..32 edit ops (replace_lines, insert_lines, replace_source, insert_cell, delete_cell, move_cell, set_cell_type, clear_outputs) | |
| path | Yes | Notebook path | |
| dry_run | No | Compute everything but do not write. Default false | |
| create_backup | No | Default true | |
| expected_content_hash | No | Optional optimistic-lock hash |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one genuinely important trait: a hash mismatch aborts the entire request atomically without partial writes. However, it says nothing about dry_run/create_backup defaults, permission needs, or concurrency scope beyond the anchor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and then the critical constraint. No filler, though the anchor sentence would be stronger if it used the real parameter names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a batch mutation tool with no annotations and no output schema, atomicity is covered well, but the description omits dry_run/backup semantics and return behavior, and it references non-existent parameters, leaving real gaps an agent must resolve from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is nominally 100% (baseline 3), but the description names 'expected_source_hash' and 'expected_text' as required anchors, neither of which exists in the schema, while the schema's actual lock parameter is 'expected_content_hash' described as merely optional. This mismatch is actively misleading about which parameter to pass and whether it is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Edit notebook cells') that cleanly separates it from siblings like notebook_read and notebook_run. It stops short of naming the alternative tools, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a precondition (a CAS anchor is required for source changes) but never says when to choose this tool over notebook_read/notebook_run or how to structure a batch edit. No when-not guidance or alternative routing is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notebook_kernelC
Inspect or manage the kernels held for notebooks: status, start, shutdown, restart.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Notebook path; required for start/shutdown/restart, ignored for status | |
| action | Yes | One of 'status' | 'start' | 'shutdown' | 'restart' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose that shutdown/restart are destructive to kernel state, whether elevated permissions are needed, or what the response looks like. Only the benign 'path ignored for status' detail is conveyed, and even that is already in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with no filler, and the resource is stated before the action list. It is appropriately brief, though for a multi-action tool the brevity edges toward under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool includes mutating actions (shutdown, restart) but has no annotations and no output schema to convey safety or return information. The description does not compensate, so an agent cannot tell how risky or reversible these operations are.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema (including path being required for start/shutdown/restart and ignored for status). The description's action list adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (kernels held for notebooks) and enumerates the operations (status, start, shutdown, restart), which is clearer than most. It does not explicitly contrast itself with siblings like notebook_run or notebook_run_status, leaving the kernel-vs-run distinction to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It lists possible actions but gives no guidance on when to choose this tool over notebook_run, notebook_run_status, or notebook_run_cancel. No prerequisites or conditions are stated; the agent must infer usage from the action names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notebook_readA
Read a Jupyter notebook: cell index, type, source and existing outputs. Set include_outputs='full' to get a cell's outputs including images.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Notebook path (absolute, or relative to the server root) | |
| cell_indexes | No | 0-based cell indexes to read; omit for all cells. This is an integer array — do not pass a range string. | |
| include_source | No | Source detail: 'none' | 'preview' (default) | 'full' | |
| include_outputs | No | Output detail: 'none' | 'summary' (default) | 'full' | |
| expected_content_hash | No | Optional optimistic-lock hash |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose the return shape and one behavioral nuance (full outputs include images), which is genuine value. However it never states that this is a non-mutating/safe operation, nor anything about cost, size limits, or behavior when the notebook is stale relative to expected_content_hash.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, front-loaded with what the tool reads before the optional flag tip. Every clause earns its place and nothing is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey the return shape, and it does list cell index, type, source and outputs. What is missing is edge behavior: whether omitted cell_indexes means all cells (only in the schema), ordering, and how truncation or errors surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents path, cell_indexes, include_source, include_outputs and expected_content_hash, making 3 the baseline. The description reinforces one parameter's effect (include_outputs='full' returns images), adding modest meaning but no syntax or default information beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (Jupyter notebook) and enumerates what is returned: cell index, type, source and existing outputs. It is easily distinguished from notebook_edit and notebook_run by the read verb, but it never names or contrasts with those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage guidance is a conditional tip: set include_outputs='full' to get outputs including images. There is no statement of when to prefer this over notebook_edit or notebook_run, nor any exclusion or prerequisite (e.g. must the kernel be idle?). Usage is implied by the read verb rather than explained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notebook_runA
Execute notebook cells. mode='resume' runs only the target cells in the live kernel; 'replay' silently rebuilds state from cell 0 first; 'full' re-runs everything.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Execution mode: 'auto' (default) | 'resume' | 'replay' | 'full' | |
| path | Yes | Notebook path | |
| cell_selector | No | Which code cells to run: 'all' (default), '3', '0-4', or a comma list like '0-4,7,9'. This is a string selector — do not pass an array. | |
| create_backup | No | Default true | |
| write_outputs | No | Write fresh outputs back to the .ipynb. Default true | |
| timeout_seconds | No | Per-cell timeout in seconds (1..86400; default from server config) | |
| clear_outputs_before | No | Clear target cells' outputs before running. Default true | |
| expected_content_hash | No | Optional optimistic-lock hash |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose one important side effect: 'replay' *silently* rebuilds state from cell 0 first, which tells the agent about hidden re-execution. It omits other key traits such as whether the call blocks or returns asynchronously (the existence of run_status/run_cancel siblings hints at async but is never stated), permission requirements, and the fact that outputs are written back to the .ipynb by default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, then a compact enumeration of the modes. Every clause carries information and nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter execution tool with no annotations and no output schema, the description covers the most important behavioral axis (kernel state per mode) but leaves the agent guessing about blocking vs async execution, timeout behavior, backup/write-back defaults, and the optimistic-lock hash. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema cannot: it explains the runtime semantics of resume/replay/full rather than just listing enum values. It leaves 'auto', cell_selector, create_backup, write_outputs, timeout_seconds, and expected_content_hash entirely to the schema, which limits it below a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Execute notebook cells') and immediately distinguishes three execution strategies by their kernel-state semantics. It does not explicitly name the sibling tools it differs from (notebook_edit, notebook_read), but the action is unambiguous enough that an agent can separate it from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mode semantics imply when each is appropriate ('resume' keeps live state, 'replay' rebuilds from cell 0, 'full' re-runs everything), which is useful implied guidance. However, it never states when to choose one over another, what 'auto' resolves to, or when to pair the call with notebook_run_status/notebook_run_cancel for long executions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notebook_run_cancelB
Cancel a background notebook run started by notebook_run.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Run id returned by notebook_run |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Cancel' implies mutation, but nothing is said about whether cancellation is graceful or forced, whether partial outputs are kept, whether the call is idempotent, or what happens for an already-completed run — all material for a destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action, with no redundant or filler content. Every word contributes to identifying the target operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema this is minimally adequate: purpose and the source of run_id are clear. However, with no annotations the description should have covered at least the effect of cancellation and its behavior on already-finished or unknown runs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents run_id as 'Run id returned by notebook_run'. The description restates this provenance but adds no format, constraint, or example beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (cancel) and resource (background notebook run), and ties the target to notebook_run as its originator, so an agent can tell it apart from notebook_run_status or notebook_run itself. It does not explicitly name a sibling it must not be confused with, which keeps it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'started by notebook_run' implies the prerequisite that a run must already exist and that only background runs are cancellable. There is no explicit guidance on checking notebook_run_status first, nor any statement of when cancellation is inappropriate (e.g. already-finished runs), so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notebook_run_statusA
Poll a background notebook run started by notebook_run.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Run id returned by notebook_run |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. 'Poll' plus 'background' usefully signals a non-blocking, repeatable status check on an async operation, but nothing is said about terminal states, error behavior, rate limits, or whether the eventual result is included.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, zero waste, with the essential information (poll, background, origin) front-loaded. Nothing is padded or repeated from the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter polling tool with no annotations and no output schema, the description is minimally adequate. It does not tell the agent what a poll returns or when the run is finished, which is the main thing an agent needs to decide whether to call again.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter with 100% schema description coverage ('Run id returned by notebook_run'), so the schema already fully documents it. The description adds no syntax or format detail beyond that, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (poll) and resource (background notebook run) and explicitly ties the run to its originator, 'started by notebook_run'. That distinguishes it from the start (notebook_run), cancel (notebook_run_cancel), and content siblings, though it never names them as alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage by stating the run must have been 'started by notebook_run', so the agent knows this is a follow-up call. However, there is no guidance on polling cadence, when to stop polling, or when to prefer notebook_run_cancel instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
notebook_edit - First observed
notebook_kernel - First observed
notebook_read - First observed
notebook_run - First observed
notebook_run_cancel - First observed
notebook_run_status
TDQS
Scored across 6 tools
Each tool targets a clearly distinct action on a notebook: read, edit, run, poll run status, cancel run, and manage kernels. The run_status/run_cancel pair are explicitly scoped to background runs started by notebook_run, so there is no real misselection risk.
All six tools use a uniform snake_case notebook_* prefix with a predictable verb/noun suffix (read, edit, run, run_status, run_cancel, kernel). The pattern is consistent throughout with no style deviations.
Six tools is well-scoped for a notebook server, covering the core read/edit/execute workflow plus async run lifecycle and kernel management. No tool feels redundant or padded.
Read, edit, run, and kernel lifecycle are covered, including async run polling and cancellation. Minor gaps remain, such as creating a brand-new notebook file or deleting cells/notebooks explicitly, but these are largely workable via edit.
Maintenance
Related MCP Connectors
- flockfsOAuthcom.flockfs
A real-time, multiplayer filesystem for agents and humans: shared files, live edits, history.
System-of-record notebook for AI coding agents: pages, datastores, tasks, skills over MCP.
Persistent cloud workspaces for AI agents: run commands, edit files, use git and a browser.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables programmatic interaction with Jupyter notebooks, allowing reading, editing, and executing cells via Claude.633MIT
- AlicenseAqualityCmaintenanceEnables AI agents to create, read, edit, and execute Jupyter notebook cells, manage kernels, and connect to remote Jupyter servers.21MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to execute Jupyter notebook cells with persistent kernel state, output persistence, and structured JSON control surface.2-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to safely and structurally edit Jupyter Notebook (.ipynb) files by providing tools to read, edit, add, and delete cells without corrupting the JSON structure.1-