marimo-inspect
A pluggable toolkit for inspecting and editing live marimo notebooks, exposed as a FastMCP server and a Python client.
Discover and bind sessions:
list_active_notebooks,set_active_session, explicitsession_id/server_urltargeting, and structured refusal reasons for missing/ambiguous/unreachable/denied targets.Read notebook state:
get_cell_map,get_cell_data,get_cell_outputs,get_variables,get_dependency_graph,get_errors, andlint_notebook.Edit the notebook:
create_cell,edit_cell(with read-before-edit staleness guard),delete_cell, andrun_cellincell,descendants, orallmodes.Interact with widgets:
set_ui_valuesets livemo.uielement values with shape validation, read-back verification, and error reporting for rejected values or failed handlers.Manage lifecycle:
restart_kernelcloses and re-materializes the kernel; session id is point-in-time.Publish static MCP resources: co-work loop, live-safety rules, and fallbacks/limitations documentation.
Python client path:
MarimoClientanddiscover_servers()for driving live sessions without requiring FastMCP.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@marimo-inspectshow me the current cells and variables in my active marimo notebook"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
marimo-inspect
A pluggable toolkit for inspecting and editing live marimo notebooks — over the wire or from inside a notebook.
It packages live-session co-work primitives as a standalone, installable component you can add to any repo:
A Python client for driving a live marimo session interactively.
A FastMCP server exposing that tooling to AI agents over the Model Context Protocol.
The two ways to use it
1. As a Python client (the "plug into any repo" path)
Install it into any project and co-work on a running marimo session from inside a notebook or script — no copy-paste of the framework:
import asyncio
from marimo_inspection import MarimoClient, discover_servers
async def main():
servers = await discover_servers()
async with MarimoClient(servers[0].url) as client:
sessions = await client.list_sessions()
result = await client.execute(sessions[0].session_id, "1 + 1")
print(result.status, result.stdout)
asyncio.run(main())Importing marimo_inspection does not require fastmcp — the MCP server
is lazily imported only when you call create_server().
2. As an MCP server (for AI agents)
from marimo_inspection import create_server
server = create_server()
server.run(transport="stdio")Or run it directly:
marimo-inspect --transport stdioFor the supported local, explicit-remote, and Tailnet shared deployment shapes,
including URL-resolution and multi-client binding rules, see
docs/deployment-topologies.md.
Related MCP server: jupyter-notebook-mcp
MCP tools
15 tools over the same live kernel.
Reads / state: list_active_notebooks, set_active_session,
get_cell_map, get_cell_data, get_cell_outputs, get_variables,
get_dependency_graph, get_errors, lint_notebook.
Writes: create_cell, edit_cell, run_cell, delete_cell.
Widget interaction: set_ui_value.
Lifecycle: restart_kernel (closes the kernel and re-materializes a fresh one;
the server process survives, but the confirmed session id is point-in-time —
session_id_stable: false). Its refused-census vocabulary is the same access
taxonomy: edit_scope_required for a readable-but-scope-denied server (run
mode), auth_required for a real auth gate, session_census_denied when the
probes cannot tell them apart.
list_active_notebooks discovers sessions and auto-binds the first one
(session_id and server_url); every other tool falls back to that
binding. The binding lives in the MCP session's server-side state plus a
process-global fallback consulted only when that state has nothing bound.
Restarting or spawning a fresh server process loses it (re-run
list_active_notebooks); over one persistent stdio process, fresh MCP
sessions still reach it through the fallback. Over HTTP the fallback is served
only while the process has seen a single client session — once a second client
session appears, argument-less calls fail closed with reason: binding_ambiguous. When in doubt, pass session_id/server_url explicitly
(or call set_active_session).
Binding is not ownership. Each session row reports
provenance: "unknown" and owner: "unknown" — marimo's public API exposes
only a session's filename/path, so who created it and which client holds it are
not knowable from here. The summary is scoped the same way: total_notebooks
is the truthful session count (always equal to session_count), session_count
counts sessions reported by GET /api/sessions, result_row_count is
len(notebooks) — every row, including a connection-failure sentinel (keys
exactly name/path/session_id/server_url/error, no
provenance/owner), which is a row but not a session. attached_client_count
is always null (marimo publishes no per-session client count, and the
server-wide /api/status/connections.active value counts sessions with an open
main consumer, not attached clients), and active_connections is a
deprecated compatibility alias for session_count that was never a client
count.
Target resolution and the common refusal
Every cell/session-targeting tool (get_cell_map, get_cell_data,
get_cell_outputs, get_variables, get_dependency_graph, get_errors,
lint_notebook, create_cell, edit_cell, run_cell, delete_cell,
set_ui_value) resolves its target through one shared step. A target problem is
therefore a structured payload, not a tool-level exception, and it is always
status: error with target_resolved: false, operation_ran: false, and
state_changed: false — no mutation is dispatched on a refusal: a
session-target refusal is built before any call at all, and a cell/anchor
reference is checked against a read-only validation of the live cells before
any write scratchpad is generated:
session_required— no explicit and no bound target (session_id/server_url); discover/bind one withlist_active_notebooksor pass both explicitly.binding_ambiguous— the bound target exists but this server process has served more than one client session, so the process-global fallback was withheld (see the binding rules above);reasonis passed through unchanged.session_not_found— the requested id is not live on the selected server; the payload'savailable_sessionsis that same server's truthful census (never another server's, never a guess) andavailable_sessions_readableistrue, because that census was actually read.server_unreachable— the session census could not be reached at all.server_query_failed— the census answered with a non-auth error (HTTP 500 included) or an unreadable body (truncated JSON or invalid UTF-8 included).edit_scope_required— the census was denied for lack of edit scope while a read-scope endpoint (GET /api/version) answered 200: the server is reachable and readable, but the census itself needseditmode. Amarimo runserver is the typical case, and authenticating cannot fix it.auth_required— the read-scope endpoint is denied with the same auth body as the census, so the server requires marimo auth even for reads.session_census_denied— the census was denied and the read-scope probes could not tell an auth gate from a missing edit scope: the cause is reported as undetermined instead of guessed.unknown_cell_ids— a cell reference (delete_cell'scell_id,edit_cell'scell_id, orcreate_cell'safter/beforeanchor) is not a live cell, by id or by name. The payload echoes the reference (anchor+anchor_cell_idforcreate_cell), lists it inunknown_cell_ids, and dispatches nothing — the raw marimoKeyErrortraceback is never returned.conflicting_anchors—create_cellwas given bothbeforeandafter.target_validation_failed— the read-only live-cell validation itself failed, so the reference could not be checked and nothing was dispatched.
A denied census cannot classify itself. Pinned marimo 0.24.0 serves an API 403
as 401 {"detail":"Authorization header required"} and strips
WWW-Authenticate, so a run-mode scope denial and a true auth gate look
identical at the denied endpoint — which is why the reason comes from a
read-scope probe (GET /api/version), never from a response header or a
byte count. All three access refusals carry read_scope_status_code and
page_kind as their evidence, and set_active_session / restart_kernel
speak the same vocabulary.
An explicit session_id/server_url pair that does not exist on the selected
server is the session_not_found case, with the real session ids in
available_sessions to correct it. Absence is concluded only from a census
that was successfully read, so an unreadable census is server_query_failed
and never session_not_found. available_sessions_readable reports which one
you got: true when the list is a real census (a successful empty census
counts), false when it is empty because nothing could be read.
For human co-work, prefer browser-first (marimo 0.24 edit mode without an
explicit --session-ttl): open the notebook so the page becomes/holds the
main consumer connection for the session, then call
list_active_notebooks and bind it — the payload owner still reads
"unknown" because the public API does not publish that role. Creating the
session yourself with the /sse handshake is for headless, agent-only work:
without an explicit --session-ttl, closing that stream leaves the session an
orphan (it outlives the stream) until a later connection takes it over — though
a configured --session-ttl can reap that orphan. A human who opens the
page must take over and re-run the notebook before its widgets respond, a
second distinct client joins the same kernel as a non-main, read-only
consumer, and a later reconnect can re-key the session id. The re-key is
separate from a page-vs-session divergence observed once: its cause is not
diagnosed, and neither the re-key nor any read-path explanation is confirmed as
its cause — so do not assume one.
Discovering a server is not discovering a session. Launching a server
creates no session in marimo 0.24 — edit and run mode alike; only a client
connect (/ws, or the browser's /sse stream) does. A fresh headless server is
therefore discoverable with an empty census, and a session listed before any
client attached is an earlier client's orphan on a still-running server (or the
browser marimo auto-opened when the launch was not headless) — never a launch
artifact, and never the on-disk __marimo__ session cache, which stores cell
outputs. A marimo run server is not discoverable at all: it registers under
--no-token like an edit server, but GET /api/sessions requires edit scope
and answers 401 in run mode, so the census-200 health check drops it and it is
never counted by a discovery-based servers_discovered. Only an explicit
server_url pointed at one reaches it, as a single connection-failure sentinel
row (session_count 0, servers_discovered 1) — and a target-resolution or
restart_kernel call against it is refused with reason: edit_scope_required
(its read-scope endpoints answer normally, so it is not an auth problem).
Read before you edit
edit_cell carries a staleness guard (check_fresh=True by default) and
requires that you read that exact cell first:
first touch of a never-read cell →
status: "needs_read"(unconditionally, even on a session with no snapshot);source changed since your last read →
status: "conflict".
Recovery is a real re-read, then retry: call get_cell_data (which records the
read baseline), then retry edit_cell. A get_cell_map preview does not
record the baseline — a preview is not a source read. A successful edit
returns the post-edit code_hash. check_fresh=False is an explicit force
escape hatch, not the recovery path. A cell id that is not live is a
structured refusal before anything is mutated (reason: unknown_cell_ids,
target_resolved/operation_ran/state_changed all false) — the same family
delete_cell and create_cell's placement anchors use.
Running cells — and running the whole notebook
run_cell(cell_id, mode=...) queues what the mode names, and reports what
each target actually did. cell_id is a cell id or cell name (resolved the
way ctx.cells resolves a key — a name that worked before the modes existed
still works): requested_cell_ids always carries the resolved IDs, cell_id
echoes your input, and resolved_cell_id reports the resolution.
mode | queues | needs |
| just that cell |
|
| that cell + its kernel-graph descendants |
|
| every document cell, in one code-mode context | an empty |
mode="all" is how an unreferenced cell (and the widgets it defines) finally
runs on a fresh session: nothing else ever pulls it in, so set_ui_value on its
widget returns unknown_variable until then. It is a full re-run — already
idle cells run again. mode="descendants" refuses rather than degrades: a
target the kernel graph has not registered (every document cell, on a fresh
un-instantiated session) returns status: error, reason: graph_unpopulated
and runs nothing. Passing mode="all" together with a cell_id is refused
(reason: cell_id_not_allowed).
The kernel may additionally run cells outside requested_cell_ids — stale
ancestors, and registered descendants in autorun mode — and the relative order of
independent cells is unspecified; requested_cell_ids is a set of targets, never
an execution order. The response is per requested target: cells[].runtime_state
and cells[].errors (with errors_readable; errors is null when the
post-run error channel could not be read, and such a target is reported
not_run/unverified, never succeeded), plus succeeded_cell_ids (idle
with a readable, empty errors), failed_cell_ids
(exception/marimo-error/cancelled/interrupted — it covers the requested
targets only) and not_run_cell_ids / unverified_cell_ids (everything else,
e.g. stale/disabled/unknown). status is ok only when every requested target
is idle, partial when any failed or did not finish, and error for
validation/planning/reporting failures (cell_id_required,
cell_id_not_allowed, unknown_cell_ids, graph_unpopulated,
planning_failed, reporting_failed). mode is a literal enum in the
published input schema, so an out-of-enum value is rejected by the framework
(literal_error) before the handler runs — it has no structured reason, and
none is documented. Every failure — and a run whose batch
call itself failed (execution_error + stderr) — carries a top-level error
string alongside the structured status/reason, so a caller that checked the
old {"error": ...} payload still sees the failure.
New cells are visible by default
create_cell defaults to hide_code=False, so a new cell's code shows in the
UI. (This changed from the earlier hidden-by-default behavior.) Pass
hide_code=True explicitly for setup/implementation cells you want hidden.
Note that no read tool echoes hide_code, so visibility is decided at creation
time and cannot be confirmed back through the MCP read surface.
Widget interaction
set_ui_value(variable_name, value) sets a live mo.ui element's value by its
kernel-global name and accepts no source code. It never coerces the value:
send the shape the element's declaration accepts — scalar for slider/text,
bool for checkbox, the option key inside a one-element list for a
dropdown (["beta"]), a list of keys (any length, [] clears it) for
multiselect, a two-element list for range_slider. A shape the element cannot
accept is refused before anything is applied: a dropdown in particular takes
exactly one key, so a list whose length is not one is refused rather than
applied (marimo would clear it to None). The error carries
reason: value_shape_mismatch and the corrected payload in did_you_mean when
one particular valid replacement can be inferred; when it cannot — an empty or
multi-element list against a dropdown with several options — did_you_mean is
absent and the message names the element's option keys instead.
The element's value is read back before the call returns, so status: ok with
verified: true means the read-back succeeded: either the widget's own value was
observed to move (applied: true) or it already held that value
(applied: false + no_change: true). A value marimo rejected — an unknown
dropdown key, say — is returned as status: error with reason: value_not_applied and the kernel's message instead of a misleading success
(a refused shape uses reason: value_shape_mismatch); a value the element's own
on_change handler raised on is reason: on_change_failed with
handler_ran: true — the value was accepted, with applied: true when it moved
or applied: false + no_change: true when the element already held it, so
only the callback failed.
A button/run_button is the exception. Both expose a click counter as
their frontend value, but their element value differs: a button's value is its
on_click return (unchanged when the handler only sets state), while a
run_button has no on_click and its value is set to True on a click and
reset to False after the dependent cells run — so either can show an unchanged
value even though the click landed. Those payloads carry the frontend counter
read (frontend_value_before/after), click_delivered, a tri-state
handler_invoked, and side_effects_verified: false. 0 is the
initialization sentinel (handler_invoked: false, no click — marimo processes
no click for it); a nonzero counter that moved to the submitted value is
handler_invoked: true; a counter that already held the submitted value is
null (unknown — the read-back cannot see a repeated click, so neither
"delivered" nor "skipped" is reported). A button whose on_click raises
returns reason: on_click_failed with handler_ran: true and acknowledges that
partial side effects may already have been applied; that marker is attributed
only to a button clicked with a nonzero counter (never a run_button or any
other element), otherwise the call fails as a generic ui_update_failed rather
than mis-attributing it.
The update is flushed and triggers reactive re-execution of dependent cells, but
set_ui_value does not verify arbitrary downstream effects: in autorun mode the
kernel re-runs dependent cells as part of the update, in lazy mode it only marks
them stale — confirm a handler's effects (including a side-effect-only button's)
with the read tools.
MCP resources
The server also publishes three static, read-only documentation resources
(text/markdown, packaged in the wheel and loaded via importlib.resources)
that a client can read on demand:
URI | Content |
| the MCP-first co-work loop, step by step |
| read-before-edit and live-kernel safety rules |
| what MCP does not cover + intentional fallbacks |
A client lists them with list_resources() and fetches one with
read_resource(uri); through the Python client you can also read them with
FastMCP's in-process transport:
from fastmcp import Client
from marimo_inspection.server import create_server
async with Client(transport=create_server()) as client:
for r in await client.list_resources():
print(r.uri, r.mime_type)
doc = await client.read_resource("workflow://marimo-inspect/live-safety")FastMCP 4.0.3 resource annotations only carry
audience/priority/lastModified, so read-only intent is carried by tags +
description; tool annotations do support readOnlyHint/destructiveHint/
idempotentHint/openWorldHint (used by set_ui_value and restart_kernel).
Output limits
marimo's code-mode snapshot exposes one main output per cell plus console
events — not every frontend UI registration. get_cell_outputs returns that
main output and the serialized console events, so a widget rendered to the user
may be absent from it; inspect the cell's variables instead.
get_cell_data returns each cell's source and live runtime_state; its
data[].variables is a deprecated compatibility placeholder (always null),
so read live variables with get_variables. To fold the error read in, pass
include_errors=True: every data[] row then also carries
structured_errors and console_stderr (the same two separate channels
get_errors reports) plus has_console_exception and
console_exception_evidence, with []/[]/false/null for a clean cell.
The default (include_errors=False) omits those four keys, so the row shape
is unchanged.
get_cell_map's has_output / has_console_output / has_errors flags are
computed from live fields (None when a private field is unreadable, never
faked). get_errors reports cells[].structured_errors (marimo cell.errors) and
cells[].console_stderr (console events, including UI-handler tracebacks) as
two separate per-cell channels; flagged entries name the matched marker in
cells[].console_exception_evidence. Top-level error flags and totals summarize
the two channels; has_errors/total_errors cover structured errors only.
lint_notebook is static and in-process (the engine behind marimo check).
Each diagnostic locates a position in the notebook source file:
rule/name/severity/message plus filename, line, column, and
cell_index — the flagged cell's positional index in the parsed document.
cell_index is not a live cell id (those come from get_cell_map); the
two are separate identifier spaces, so resolve the live id separately and
navigate the source by line/column/filename.
Arbitrary kernel probes, complex multi-operation CodeMode blocks, screenshots,
and notebook-server lifecycle stay outside the MCP surface — see
reference://marimo-inspect/fallbacks-and-limits.
Install and connect an MCP client
The MCP resources explain how to operate a connected live notebook; they cannot bootstrap their own installation. Start here, then use the packaged resources once the harness reports the server connected.
Normal consumer installation
Add a pinned release to the notebook project. This is the standard path for users and consumer-repository contributors; it is a normal, non-editable installation in that project's environment.
uv add "marimo-inspect @ git+https://github.com/ajegorovs/marimo-mcp-cowork@v0.3.3"
uv syncConfigure the harness to execute that environment's console script, not
uv run and not a provider checkout:
<project-root>/.venv/bin/marimo-inspect --transport stdioUse the relevant configuration block in
docs/harness-integration/README.md, then
verify the installed script with .venv/bin/marimo-inspect --help. Start a
marimo notebook with --no-token, open it in a browser to materialize a
session, and call list_active_notebooks. After connection, read the MCP
resources for the co-work loop and safety rules.
Provider contributors: local override only
Use an editable sibling checkout only when testing unreleased changes to this provider against a consumer project. It is not a consumer-installation mode: replace the consumer's pinned dependency temporarily, resync, test, then restore the pinned version. Do not commit a machine-local dependency source.
Examples
examples/ holds consumer-facing marimo usage patterns:
notebooks that run on marimo alone, with none of this package's code or
dependencies. They are patterns to copy, not fixtures — nothing under
examples/ is imported by tests and the package never needs them.
Example | Pattern |
Two dependent controls (the child's options come from the parent's value) hosted in the sidebar. | |
A slider paired with step buttons over one shared |
Check them statically (the full notebook check covers both trees):
uv run marimo check notebooks examplesEach example is also booted on a real kernel and executed whole by the live
suite (tests/marimo_inspect/live/test_examples.py, run with -m live), which
additionally drives their live controls. See
examples/README.md for the pattern notes and the
verification split.
Requirements
Python 3.12+
A running marimo edit server (start it with
--no-tokenfor registry-based discovery; seediscover_servers— amarimo runserver registers but is not discoverable, because its session census requires edit scope).marimo 0.24.x (private APIs are version-bound — see docs/marimo-version-support.md for the pinned range and the upgrade validation procedure).
Development
uv sync --all-extras # test deps are an optional extra; a bare sync prunes pytest
uv run ruff check .
uv run pytest -m "not live" # unit tests (fast, no kernel)
uv run pytest -m live # live kernel tests (boots its own headless server)The live tests need a real marimo kernel and are deselected by default. See
docs/live-tests.md for how the live suite is run and its
current status.
Available Tools
14 toolscreate_cellCreate CellB
Create a new cell in the notebook.
Created cells are visible in the UI by default (hide_code=False); pass hide_code=True explicitly for setup/implementation cells you want hidden.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Optional cell name. | |
| after | No | Optional cell_id to place this cell after. | |
| before | No | Optional cell_id to place this cell before. | |
| source | Yes | Source code for the new cell. | |
| hide_code | No | Whether the code is hidden in the UI (default False). | |
| server_url | No | Server URL override. | |
| session_id | No | Session ID; omit only when the active-session binding holds for this call (see `list_active_notebooks`). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses a default outcome ('cells are visible in the UI by default'), but says nothing about session binding/auth requirements, whether the cell is persisted or executed, or what occurs if both after and before are supplied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the default-behavior point is front-loaded. The opening sentence largely restates the name/title, which is conventional but not additive, keeping it just below a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the schema fully covers parameters. However, for a creation tool with no annotations, the description omits permission/session prerequisites and the interaction between the mutually related after/before placement options, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters are already documented in the schema; this sets the baseline at 3. The description adds a modest amount of meaning for hide_code by explaining the intent behind the flag, but nothing for name, after, before, server_url, or session_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a new cell in the notebook'), which cleanly separates it from siblings like edit_cell, delete_cell, and run_cell. It does not explicitly name those siblings or clarify boundaries, but the create/edit/delete/run distinction is unambiguous from the verb alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real guidance for one decision point: hide_code defaults to False, and True is for 'setup/implementation cells you want hidden.' That is usage guidance for a parameter rather than for the tool itself; there is no guidance on when to use create_cell versus edit_cell, or on the after/before positioning options.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_cellDelete CellC
Delete an existing cell from the notebook.
| Name | Required | Description | Default |
|---|---|---|---|
| cell_id | Yes | Target cell id. | |
| server_url | No | Server URL override. | |
| session_id | No | Session ID; omit only when the active-session binding holds for this call (see `list_active_notebooks`). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. 'Delete' implies a destructive mutation, but the description does not state whether deletion is permanent, reversible, requires specific permissions, or affects dependent cells. It adds little beyond the operation name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no wasted words and the action is front-loaded. It is appropriately concise, though for a destructive operation it could benefit from one more sentence of context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a destructive tool with no annotations, no stated reversibility or permissions, and alternatives among siblings, the description is insufficiently complete. While the output schema exists and parameters are fully documented in the schema, the lack of any behavioral or safety context leaves an agent without critical information for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all three parameters (cell_id, server_url, session_id) with descriptions. The tool description adds no additional meaning about parameter usage or formats. The baseline of 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete an existing cell from the notebook.' It is clear what the tool does. However, it does not explicitly differentiate itself from siblings like edit_cell or create_cell beyond the obvious verb, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as edit_cell or create_cell. There are no exclusions, prerequisites, or contextual triggers mentioned. It only states the action, leaving usage entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_cellEdit CellA
Edit an existing cell's source code.
Includes a staleness guard: unless check_fresh=False, refuses to edit
a cell that has no full-source read baseline (needs_read) or whose
source changed since that read (conflict). get_cell_data records
the baseline; a get_cell_map preview does not. This prevents silently
overwriting a concurrent edit.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Optional new cell name. | |
| source | Yes | New source code. | |
| cell_id | Yes | Target cell id. | |
| hide_code | No | Optional new hide_code value. | |
| server_url | No | Server URL override. | |
| session_id | No | Session ID; omit only when the active-session binding holds for this call (see `list_active_notebooks`). | |
| check_fresh | No | Refuse to edit a cell that changed since last read. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does it well: it discloses the staleness guard, the two refusal conditions, and the purpose (preventing silent overwrite of a concurrent edit). It stops short of describing permission/auth requirements or what a successful edit returns, but the guard disclosure is substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the guard explained in two tight sentences with no filler. The parenthetical callout of needs_read/conflict is economical and information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 7 params, an output schema, and full schema coverage, the description supplies the critical non-obvious behavior (the freshness guard) that the schema alone would hide. Minor gaps remain around session requirements and success semantics, but they are largely covered by structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning for check_fresh beyond the schema's one-liner by spelling out the two staleness failure modes. The remaining parameters (name, hide_code, source, session_id) are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edit an existing cell's source code') and implicitly separates itself from create_cell/delete_cell/run_cell. The guard discussion further delineates it by naming get_cell_data and get_cell_map as related-but-different operations, so an agent can place it precisely among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when the edit proceeds and when it is refused (needs_read with no read baseline, conflict when source changed), plus the check_fresh=False escape hatch. It also tells the agent how to satisfy the precondition (use get_cell_data, not a get_cell_map preview), which is genuine routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cell_dataGet Cell DataA
Get full runtime data for one or more cells.
Includes source code, errors, and variable information. If cell_ids is empty, returns data for all cells.
A requested id that resolves to nothing (deleted or mistyped) is reported
in missing_cell_ids rather than silently omitted — the write tools
refuse an absent id, so an unqualified empty payload here would hide the
same mistake.
| Name | Required | Description | Default |
|---|---|---|---|
| cell_ids | No | Cell IDs from get_cell_map. Empty = all cells. Accepts a single ID, a native array, or a JSON-encoded array — a harness may deliver either of the latter two as a string. | |
| server_url | No | Optional server URL override. | |
| session_id | No | Session ID from list_active_notebooks. Optional if an active session is bound. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose genuinely non-obvious behavior: empty cell_ids returns all cells, and unresolvable ids surface in missing_cell_ids instead of being silently dropped. The rationale (write tools refuse absent ids, so silence would hide the mistake) adds real context. It omits cost/pagination implications of fetching all cells, so it is not fully complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in sentence one, followed by scope, then the edge-case semantics. The final explanatory clause is slightly long but earns its place by justifying the missing_cell_ids design. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't explain return shapes, and it covers the default/all-cells case plus the error-reporting edge case. The remaining gap is disambiguation from the closely related read siblings, which a three-parameter read tool in this family would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents cell_ids semantics, the accepted string/array/null shapes, server_url, and session_id. The description restates the empty-equals-all behavior without adding new format or constraint detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first line gives a specific verb and resource ("Get full runtime data for one or more cells") and enumerates what 'runtime data' means (source code, errors, variables). It does not, however, distinguish itself from siblings that overlap heavily with that content — get_errors, get_variables, and get_cell_outputs are adjacent tools and the description never says how this differs from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the empty-cell_ids behavior tells the agent this can be used as a bulk fetch, but there is no statement of when to prefer this over get_cell_map, get_errors, or get_variables. No exclusions, prerequisites, or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cell_mapGet Cell MapA
Get a lightweight map of cells showing previews.
Returns cell IDs, names, code previews, and runtime state. This is the starting point for navigating a notebook.
A preview does NOT record the edit_cell read baseline — it updates only
this tool's own change-detection snapshot. Read a cell's full source with
get_cell_data before editing it.
| Name | Required | Description | Default |
|---|---|---|---|
| server_url | No | Optional server URL override. | |
| session_id | No | Session ID from list_active_notebooks. Optional if an active session is bound. | |
| preview_lines | No | Number of lines to show per cell (default: 3). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses a genuinely non-obvious behavior: a preview does not record the edit_cell read baseline and only updates this tool's own change-detection snapshot. That is valuable. It still omits other behavioral context such as how the snapshot interacts across sessions or whether the map is cached, so it falls just short of full coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return contents in two short sentences, followed by one earned cautionary paragraph about the edit baseline. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be explained, yet the description still summarizes it. Combined with the preview/edit-baseline caveat and clear routing to get_cell_data, an agent has everything needed to call this safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so server_url, session_id, and preview_lines are already documented in the schema. The description adds no extra semantics about these parameters (e.g., preview_lines tradeoffs), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (get a map of cells) and states exactly what it returns: cell IDs, names, code previews, and runtime state. It also positions itself against the sibling get_cell_data by calling itself the 'starting point for navigating a notebook' and contrasting previews with full source.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use it ('starting point for navigating a notebook') and names the alternative with a condition: read full source with get_cell_data before editing. The agent does not need to infer the routing decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cell_outputsGet Cell OutputsB
Get cell execution outputs including visual display and console streams.
A requested id that resolves to nothing (deleted or mistyped) is reported
in missing_cell_ids rather than silently omitted — a cell with no
output and a cell that does not exist must not look alike.
| Name | Required | Description | Default |
|---|---|---|---|
| cell_ids | No | Cell IDs from get_cell_map. Empty = all cells. Accepts a single ID, a native array, or a JSON-encoded array — a harness may deliver either of the latter two as a string. | |
| server_url | No | Optional server URL override. | |
| session_id | No | Session ID from list_active_notebooks. Optional if an active session is bound. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It does disclose one genuinely useful non-obvious behavior — unresolved IDs surface in missing_cell_ids so an empty cell is distinguishable from a nonexistent one — but says nothing about output size limits, truncation, pagination, or whether retrieval has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loading the core purpose before the edge-case note, with no filler. The second sentence is slightly roundabout in phrasing but every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and all three parameters are schema-documented. The missing-id behavior is covered; the only gap is guidance on selecting this tool over sibling readers and any limits on output volume.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so cell_ids, server_url, and session_id are already fully documented in the schema. The description adds no semantics on top of that (it does not even restate the empty-means-all-cells default), so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (get cell execution outputs) and enumerates what those outputs contain (visual display, console streams). It does not name or contrast with the nearest siblings (get_errors, get_cell_data, run_cell), so an agent must infer the boundary itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what is retrievable but never states when to call this instead of get_cell_data or get_errors, nor any prerequisite such as having run the cell first. Usage is only implied by the resource name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dependency_graphGet Dependency GraphB
Get the cell dependency graph showing variable relationships.
Reveals which variables each cell defines and references, parent/child relationships between cells, variable ownership, and dependency issues like multiply-defined variables or cycles. The graph is always the FULL notebook graph.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | NOT IMPLEMENTED — supplying a non-zero value is refused for the same reason. | |
| cell_id | No | NOT IMPLEMENTED — supplying it is refused (``reason: unsupported_argument``) instead of silently ignored. | |
| server_url | No | Optional server URL override. Optional if an active server_url is bound. | |
| session_id | No | Session ID from list_active_notebooks. Optional if an active session is bound. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, and it does disclose one meaningful trait: the graph is always full-notebook and cannot be scoped. It says nothing about permissions, cost/latency on large notebooks, or failure modes, leaving real gaps for a graph-traversal read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose, and the trailing 'always FULL notebook graph' caveat is placed after the content list where it belongs. Minor cost: the second sentence is a dense enumeration that could be trimmed without losing much.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and all four parameters are documented in the schema. What remains missing is the usage layer: no stated preconditions (active session/server_url binding) and no guidance on choosing this tool over get_variables, which is the main decision an agent faces here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the two non-trivial parameters (depth, cell_id) already carry explicit NOT IMPLEMENTED refusal notes in the schema. The description adds no parameter-level meaning, so the baseline 3 for high schema coverage applies even though the description's 'always FULL graph' line loosely reinforces why those args are refused.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb and resource ('Get the cell dependency graph') and the second enumerates the actual content (variable relationships, parent/child edges, ownership, cycles). That materially separates it from get_variables or get_cell_map, but no sibling is named explicitly, so an agent must infer the boundary from the content list alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use statement, no prerequisite (e.g. active session required), and no routing to alternatives such as get_variables for a flat variable list. The only usable cue is the implicit 'always the FULL notebook graph' constraint, which hints at when not to bother with scoping rather than when to call the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_errorsGet ErrorsA
Get all errors in the notebook session, organized by cell.
Two channels are reported and never conflated:
structured_errors: marimo's structured CellError records (kind graph|runtime, msg, exception).console_stderr: serialized stderr console events (same shape as get_cell_outputs), so UI-handler exception tracebacks are visible even when the structured channel is empty.
has_errors / total_errors / total_cells_with_errors are the
STRUCTURED-only counts (backward-compatible); has_console_exception /
total_console_exception_cells cover the console channel.
| Name | Required | Description | Default |
|---|---|---|---|
| server_url | No | Optional server URL override. Optional if an active server_url is bound. | |
| session_id | No | Session ID from list_active_notebooks. Optional if an active session is bound. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full disclosure and does so well: it explains that two channels (structured_errors, console_stderr) are never conflated, that stderr tracebacks surface even when the structured channel is empty, and that has_errors/total_errors are structured-only backward-compatible counts. This is meaningful behavioral context beyond structured fields. It stops short of describing side effects or cost, though for a read-only diagnostic tool that gap is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence front-loads the purpose, and the channel breakdown is organized and to the point. It is somewhat dense with parenthetical field lists, but every clause conveys distinct semantics rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description still clarifies the non-obvious semantics of the count fields and channel separation. The main omission is guidance on how this differs from lint_notebook when an agent is troubleshooting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so server_url and session_id are already documented in the schema, including their optional-if-bound behavior. The description adds no parameter meaning, which is acceptable at full coverage but not additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get all errors in the notebook session, organized by cell') and immediately scopes what is returned. It does not explicitly differentiate itself from the sibling lint_notebook, leaving the agent to infer the static-vs-runtime distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to reach for this tool versus alternatives such as lint_notebook or get_cell_outputs. It references get_cell_outputs only to describe a shared output shape, not to route the agent, and offers no prerequisites or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_variablesGet VariablesA
Get tables and variables information in the session.
Returns information about kernel variables and DataFrames. If variable_names is empty, returns all variables — meaning the notebook's own session names, with the inspection template's scaffolding (its imports and helpers) excluded, since the scratchpad shares the kernel namespace.
| Name | Required | Description | Default |
|---|---|---|---|
| server_url | No | Optional server URL override. | |
| session_id | No | Session ID from list_active_notebooks. Optional if an active session is bound. | |
| variable_names | No | Specific variables to inspect. Empty = all. Accepts a single name, a native array, or a JSON-encoded array — a harness may deliver either of the latter two as a string. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does disclose a genuinely non-obvious trait: empty variable_names returns the notebook's own session names with the inspection template's scaffolding excluded, since the scratchpad shares the kernel namespace. 'Get' implies a read-only operation, and return values are covered by the output schema, so the remaining gap is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose and then the empty-argument semantics. The parenthetical scaffolding explanation earns its place, though the phrasing is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be described, and the read-only nature is implied. The description covers purpose, scoping behavior, and the empty-argument case, leaving little an agent needs that is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, establishing a baseline of 3, and the description adds meaning beyond the schema by explaining what 'empty' actually returns — the notebook's own names minus template scaffolding — which the schema's bare 'Empty = all' does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'Get tables and variables information in the session,' and clarifies it returns kernel variables and DataFrames. It is clearly distinguishable from sibling cell/error tools, though it never explicitly names an alternative for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by context ('in the session', inspecting kernel state), but there is no explicit when-to-use guidance and no exclusions relative to siblings like get_cell_outputs or get_dependency_graph. An agent can infer intent but is not directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lint_notebookLint NotebookB
Lint a marimo notebook to check for issues.
Uses marimo's internal linting engine (the same one behind
marimo check) to check for:
Breaking issues: Problems that prevent the notebook from running
Runtime issues: Problems that may cause unexpected behavior
Formatting issues: Code style and formatting problems
| Name | Required | Description | Default |
|---|---|---|---|
| server_url | No | Optional server URL override. Optional if an active server_url is bound. | |
| session_id | No | Session ID from list_active_notebooks. Optional if an active session is bound. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full load. It usefully discloses that this is the same engine behind 'marimo check' and enumerates breaking/runtime/formatting categories, which conveys the nature of the output. But it never states whether the tool only reports or also mutates the notebook, nor that an active session/server is required — relevant given sibling set_active_session and the optional server_url/session_id params.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action in the first sentence, followed by a compact categorized bullet list. No filler, though the 'marimo check' aside is slightly redundant with the engine statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the three issue categories give a good sense of what comes back. What's missing is the read-vs-write nature of the operation and session requirements, which the schema only partly covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both optional params (server_url, session_id) are already fully documented in the schema, including their fallback to bound values. Baseline 3 applies; the description adds nothing about parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Lint a marimo notebook') and enumerates the three issue classes it detects. It does not, however, differentiate itself from the sibling get_errors, which an agent might reasonably confuse with linting output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to lint versus when to call get_errors, run_cell, or any other diagnostic sibling, and no stated preconditions. The 'marimo check' reference gives provenance but not a usage condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_active_notebooksList Active NotebooksA
List currently active marimo notebooks.
Returns all active sessions with their file paths and session IDs.
The first discovered session is automatically bound as the active
session, so later calls that share this MCP session can omit both
session_id and server_url.
The binding is server-side state keyed by the MCP session identity the
client negotiates, so it reaches the next call only when the client keeps
one MCP session for the connection — an mcp-SDK-based client does,
fastmcp's own Client does not on the pinned fastmcp 4.0.3 (it starts a
new session per request). Where that is the case, pass session_id and
server_url explicitly on every call.
| Name | Required | Description | Default |
|---|---|---|---|
| server_url | No | Optional explicit server URL. If not provided, discovers servers from the marimo registry. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and meets it: it discloses server-side binding state, how that state is keyed (MCP session identity), and the non-obvious client-dependent caveat that makes the binding unreliable. This is exactly the kind of stateful side effect an agent could not infer from the name or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in a single short sentence, followed by return contents and then the binding caveat. The third paragraph is dense and version-specific, but each sentence carries load-bearing information about when the binding fails, so the length is justified rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description briefly summarizes them (file paths and session IDs). Combined with the binding semantics and the client caveat, an agent has everything needed to call this correctly and to know what happens on the next call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is only one optional parameter, so the baseline is 3. The description earns an extra point by explaining the behavioral consequence of omitting `server_url` — it ties omission to the auto-binding of the first discovered session, which the schema's "discovers servers from the marimo registry" note does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb and resource ("List currently active marimo notebooks") and is the only discovery/listing tool among siblings dominated by cell/notebook mutation and inspection tools. An agent can distinguish it immediately without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the calling pattern explicitly: the first discovered session is auto-bound, so subsequent calls may omit `session_id` and `server_url`. It goes further and names the condition under which that optimization fails (fastmcp's own Client on pinned fastmcp 4.0.3 negotiating a new MCP session per request), telling the agent exactly when to pass both parameters explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_cellRun CellC
Run (execute) an existing cell.
| Name | Required | Description | Default |
|---|---|---|---|
| cell_id | Yes | Target cell id. | |
| server_url | No | Server URL override. | |
| session_id | No | Session ID; omit only when the active-session binding holds for this call (see `list_active_notebooks`). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says nothing about side effects (state mutation in the notebook, execution ordering, whether it blocks), authentication needs, or the session-binding requirement hinted at in the schema. 'Run' implies execution but no consequences are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no waste and the key action front-loaded. It is efficient, though arguably too terse given the tool's execution side effects.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a state-mutating execution tool with zero annotations the description is inadequate. It omits side effects, session requirements, and any interaction with output-retrieval siblings, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (cell_id, server_url, session_id) are already documented in the schema. The description adds nothing beyond the schema, which is the baseline 3 for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a clear verb+resource ('Run (execute) an existing cell'), which distinguishes it from create_cell, edit_cell, and delete_cell by action. However, it does not differentiate it from other read/observation siblings like get_cell_outputs or get_cell_data, leaving the scope slightly ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no mention that running a cell is the prerequisite for retrieving outputs via get_cell_outputs, and no indication of session prerequisites. The agent must infer all usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_active_sessionSet Active SessionA
Set the active notebook session for subsequent tool calls.
Binds a session_id (and optionally its server_url) so later calls can omit
both session_id and server_url. The binding is written to two places:
this call's MCP-session state (visible to a client that keeps one MCP
session across calls — an mcp-SDK-based client does) and a
process-global fallback consulted only when the MCP-session state has
nothing bound. Over stdio one process serves exactly one client, so the
fallback carries the binding to every later call, fastmcp's own Client
(a fresh MCP session per request on the pinned fastmcp 4.0.3) included.
Over HTTP/SSE one process serves many clients, so the fallback is served
only while the process has seen a single client session; once a second
distinct client session appears, argument-less calls are refused with
reason: binding_ambiguous rather than risk handing one client another
client's notebook — pass both arguments explicitly there. Use
list_active_notebooks first to discover available sessions and their IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| server_url | No | Optional server URL to bind with the session. | |
| session_id | Yes | The session ID to set as active. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses where the binding is written (MCP-session state plus process-global fallback), how stdio vs HTTP/SSE differ, and that ambiguous argument-less calls are refused with reason: binding_ambiguous rather than leaking another client's notebook. That is exactly the non-obvious behavior an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence and every subsequent sentence carries behavioral detail, but the middle block is a dense wall of text mixing stdio/HTTP semantics that could be split for faster scanning. Not wasteful, just heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described. Given the tool's deceptively stateful nature and lack of annotations, the description covers the persistence model, fallback scope, and failure mode thoroughly — nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema: that session_id and server_url become omittable in later calls because the binding persists, and why server_url is optionally bound alongside. This exceeds the schema's per-field text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set the active notebook session') and scopes the effect to 'subsequent tool calls', which cleanly separates it from siblings like list_active_notebooks or get_cell_map. An agent can tell what state this mutates without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to list_active_notebooks first to discover sessions and their IDs, and tells it to pass both arguments explicitly in the multi-client HTTP/SSE case. It lacks an explicit 'when not to use' framing, but the conditional guidance is unusually concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_ui_valueSet Ui ValueADestructive
Set the value of a live marimo UI element, by its variable name.
The kernel global named by variable_name must resolve to a marimo UI
element (e.g. mo.ui.slider, mo.ui.dropdown, mo.ui.text). Its
value is replaced with value, triggering reactive re-execution the
same way a user interaction would.
Value shape is per widget and is NEVER coerced: a slider/text takes a
scalar, a dropdown takes its option key inside a one-element list (for
example ["beta"]), a multiselect takes the list of selected keys, a
range_slider takes a two-element list, a checkbox takes a bool. The
element's own declaration decides; a shape mismatch is refused before
anything is applied and the response carries the corrected payload in
did_you_mean.
This tool accepts NO source code: it exists for widget interaction only,
not for arbitrary code execution. The update is flushed on code-mode
context exit, the kernel then re-runs dependent cells, and the element's
value is read back before returning — status: ok with
verified: true means the element's own value was observed to move (or
was already equal), not merely that the update was queued. A value marimo
rejected is reported as an error, never as success.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | New value for the element, in the shape that element accepts. | |
| server_url | No | Server URL override. | |
| session_id | No | Session ID; omit only when the active-session binding holds for this call (see `list_active_notebooks`). | |
| variable_name | Yes | Name of the live kernel global holding the UI element. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations: the update triggers reactive re-execution like a user interaction, is flushed on code-mode context exit, and values are verified by read-back so 'status: ok' with 'verified: true' means the value actually moved. It also discloses the refusal path (shape mismatch rejected before applying, corrected payload in did_you_mean) and that rejected values are never reported as success.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then layered detail in descending priority (type constraint, value shapes, then execution/verification semantics). It is dense and multi-paragraph but nearly every sentence carries non-redundant information; the value-shape enumeration is long but earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with a live output schema, it covers everything an agent needs: preconditions, value shapes, error/refusal behavior, timing of the flush, and the meaning of the verification fields. The session-binding caveat is delegated to the session_id schema and list_active_notebooks, which is coherent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real per-parameter meaning: 'value' shape is enumerated per widget type (scalar for slider/text, one-element list for dropdown, list of keys for multiselect, two-element list for range_slider, bool for checkbox) and is never coerced. The server_url/session_id parameters are left to the schema, which already documents them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'Set the value of a live marimo UI element, by its variable name.' It immediately distinguishes itself from the cell-editing siblings by stating it accepts no source code, so an agent can tell it apart from edit_cell/run_cell without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear precondition (the kernel global must resolve to a marimo UI element, with examples of types) and an explicit when-not ('NO source code … not for arbitrary code execution'), which routes the agent away from misuse. It does not name a sibling alternative for the code-execution case, so it stops short of full when/when-not/alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v0.3.3- First observed
create_cell - First observed
delete_cell - First observed
edit_cell - First observed
get_cell_data - First observed
get_cell_map - First observed
get_cell_outputs - First observed
get_dependency_graph - First observed
get_errors - First observed
get_variables - First observed
lint_notebook - First observed
list_active_notebooks - First observed
run_cell - First observed
set_active_session - First observed
set_ui_value
TDQS
Scored across 14 tools
The 14 tools generally have distinct purposes, but get_cell_map, get_cell_data, and get_cell_outputs all involve retrieving cell information and could be confused by an agent unfamiliar with the subtle differences. The descriptions do clarify the differences (preview vs full data vs outputs), so it's mostly distinct.
All tool names follow a consistent snake_case verb_noun pattern (get_*, list_*, create_*, edit_*, run_*, delete_*, set_*), with no deviations. The naming is predictable and readable.
The server has 14 tools, which is well within the typical 3-15 range for a well-scoped MCP server. Each tool appears to earn its place by covering a specific operation in notebook inspection and manipulation.
The toolset covers a broad range of notebook operations (cell CRUD, execution, outputs, errors, linting, variables, dependency graph, UI interaction, session management), but some potential gaps exist, such as creating or deleting notebooks, managing files, or deeper kernel operations. It is largely complete for inspection and basic editing.
Maintenance
Related MCP Connectors
Live browser debugging for AI assistants — DOM, console, network via MCP.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Create, edit, preview, and deploy full-stack web apps to a live URL from any MCP client.
Create, browse, remix, collaborate on, and run durable AI workflow nodes from MCP hosts.
Related MCP Servers
- FlicenseAqualityDmaintenanceAuto-discovers running marimo notebooks and exposes tools for reading, editing, and running cells via MCP, supporting both HTTP and VS Code backends.10-
- FlicenseBqualityDmaintenanceA FastMCP server for loading, editing, searching, and saving Jupyter notebooks (.ipynb) through MCP tools. It maintains a single active notebook session with live cell indices that update as cells are inserted or removed.10-
- FlicenseAqualityCmaintenanceMCP server for structural editing of Jupyter notebook cells (list, read, insert, edit, patch, delete, move) without kernel execution.10-
- FlicenseNot gradedqualityDmaintenanceEnables live runtime inspection of any Python application, allowing MCP clients to query state, evaluate expressions, inspect objects, and read source code while the app runs.-