Skip to main content
Glama

OpenReadout

OpenReadout is an open-source reader for lab-instrument files, designed for AI agents.

Give AI agents full access to raw data from microscopes, mass spectrometers, cytometers, electrophysiology rigs, and 90+ other instrument file formats — in one command.

Open-source. Single binary. No vendor software. No dependencies. No network access. Works everywhere.

OpenReadout makes data stored in proprietary instrument file formats readable: it pulls out the metadata, images, traces, spectra, and tables as structured JSON and renders previews so your agent can see and understand the data. Every format is validated against real data and independent libraries.

CI Docs License: MIT OR Apache-2.0

For AI Agents — Get Started in One Line

Paste this into your AI agent's chat — it will read the skill file and install everything:

curl -fsSL https://raw.githubusercontent.com/openreadout/openreadout/main/skills/openreadout/SKILL.md

That's it. The skill file tells the agent how to install the binary and how to use every command.

Related MCP server: Xberg MCP Server

For Humans

Option A — Browser: Open the browser demo and drop a file on it. It runs OpenReadout compiled to WebAssembly inside the page; nothing is uploaded.

Option B — CLI: Install the binary (see Installation), then connect it to your agent:

openreadout self skill --install all      # skill for Claude Code, Codex, Cursor, Copilot, Gemini CLI
openreadout mcp --install claude-desktop  # MCP server for Claude Desktop (or cursor, codex, vscode, ...)

Your agent can now open, check, plot, and convert instrument files on your behalf.

For Developers — See It Live in 30 Seconds

# 1. Install (until the first release, build from source)
cargo install --locked --git https://github.com/openreadout/openreadout openreadout

# 2. See what is in a file — reads headers only, fast on any size
openreadout info cells.lif

# 3. Look at it — writes cells.preview.png
openreadout preview cells.lif --composite

# 4. Convert it — read back and verified before it is saved
openreadout export cells.lif -o cells.ome.tiff

That's it. The same commands work on a CZI, an ND2, a Thermo RAW, an ABF, or any of the other formats.

Quick Start

# What is in the file?
openreadout info cells.lif
# → format: Leica LIF (lif) v2  size: 16.0 MiB  images: 1  planes: 2
# →   [0] PEI_laminin_35k  2048x2048 z=1 c=2 t=1  uint16  px=0.3250 µm
# →       objective: HC PL FLUOTAR L 20x/0.40 DRY

# Is it complete?
openreadout check partial-copy.lif
# → error  truncated       block chain runs past end of file
# → error  missing_planes  geometry needs 16777216 bytes but only 8969789 are stored

# Integrate the peaks of a chromatogram
openreadout analyze peaks gc-run.ch --min-height 1
# → 4 peaks, area in pA·min
# → 1   4.852 min  area 0.2779  25.08 %
# → ...

# Structured JSON for scripts and agents
openreadout info cells.lif --json
{
  "ok": true,
  "schema_version": "1",
  "data": {
    "format": { "id": "lif", "name": "Leica LIF", "vendor": "Leica Microsystems" },
    "images": [
      {
        "size_x": 2048, "size_y": 2048, "size_c": 2,
        "pixel_type": "uint16",
        "physical_size": { "x": 0.325, "y": 0.325, "unit": "µm" }
      }
    ]
  }
}

Why OpenReadout?

What used to take vendor software or a different library for every format:

import czifile, nd2, liffile, pyabf, flowio
# ... a different API, metadata layout, and set of quirks for each one ...

Now takes one command, for all of them:

openreadout info any-file --json

What OpenReadout can do:

  • Inspect images, channels, traces, spectra, tables, and metadata -- in plain text or structured JSON

  • Check files for truncation, missing planes, and damaged structure -- exit code 4 when a file is corrupt

  • Export to OME-TIFF, OME-Zarr, mzML, NWB, CSV, Parquet, Arrow, JCAMP-DX, Allotrope ASM, and RDML -- every export read back and verified

  • Preview image planes, traces, spectra, and plate heat maps as PNG

  • Analyze chromatographic peaks, plate assays (IC50, standard curves), qPCR (Cq, ΔΔCq), NMR peaks, patch-clamp features, spikes, and flow-cytometry gates -- with documented methods

  • Batch over whole directories, index lab shares, and watch running acquisitions

Area

Formats

Export to

Light microscopy

Zeiss CZI, Nikon ND2, Leica LIF, Olympus OIR/VSI/OIB, Imaris, OME-TIFF and other TIFF variants, OME-Zarr, whole-slide images

OME-TIFF, OME-Zarr

High-content screening

Harmony (Opera Phenix, Operetta), ImageXpress, CellVoyager

OME-Zarr plate, OME-TIFF

Electron microscopy

MRC, Gatan DM3/DM4, FEI SER/EMI, Velox EMD

OME-TIFF, OME-Zarr

Mass spectrometry

Thermo RAW, Bruker timsTOF, Agilent MassHunter, Waters MassLynx, Sciex WIFF, mzML

mzML, Parquet, Arrow

Chromatography

Agilent ChemStation and OpenLab, Shimadzu, Chromeleon, AIA/ANDI

CSV, JCAMP-DX, Parquet

Electrophysiology

Axon ABF, Intan, SpikeGLX, Open Ephys, Neuralynx, Blackrock, Plexon, HEKA, Spike2, NWB

NWB, CSV, Parquet

NMR and spectroscopy

Bruker TopSpin and OPUS, Varian, JEOL, Thermo OMNIC, Renishaw, JCAMP-DX, SPC

JCAMP-DX, CSV

Flow cytometry

FCS, FlowJo workspaces, Gating-ML

CSV, Parquet, Arrow

Plate readers and qPCR

Plate-reader exports, RDML, Applied Biosystems, LightCycler, Rotor-Gene

Allotrope ASM, RDML, CSV

Other

ÄKTA, ITC, Biacore, Seahorse, Octet, Zetasizer, XRD, EPR, electrochemistry, thermal analysis

CSV, Parquet

The format list has all 96 formats and their known gaps.

Use Cases

For Researchers:

  • Open instrument files on any computer, without the acquisition software

  • Convert a folder of raw files to OME-Zarr, mzML, or NWB for analysis and sharing

  • Verify that files copied off an instrument PC are complete

For AI Agents:

  • Answer questions about a file: channels, pixel size, objective, acquisition time, scan count

  • Extract metadata, traces, spectra, and tables as JSON

  • Run documented analyses (peak areas, IC50s, Cq values) and report the method used

For Core Facilities and Pipelines:

  • Index a lab share into searchable Parquet tables with index and search

  • Watch instrument directories and flag stalled or damaged acquisitions with watch

  • Run in Nextflow, Snakemake, and Galaxy pipelines (integrations/)

Installation

Ships as a single self-contained binary. No Java, no Python, no vendor DLLs -- nothing else to install.

OpenReadout has not had its first release yet. Until then, build from source with Rust 1.91 or newer:

cargo install --locked --git https://github.com/openreadout/openreadout openreadout

From the first release on:

# macOS / Linux
curl -fsSL https://raw.githubusercontent.com/openreadout/openreadout/main/scripts/install.sh | sh

# Windows (PowerShell)
irm https://raw.githubusercontent.com/openreadout/openreadout/main/scripts/install.ps1 | iex

# Homebrew (macOS / Linux)
brew install openreadout/tap/openreadout

# npm (all platforms — fetches the native binary for your platform)
npm install -g openreadout

Docker, Nix, cargo-binstall, and the other channels are on the install page.

Verify installation: openreadout --version

AI Integration

MCP Server

Built-in MCP server — register with one command:

openreadout mcp --install claude          # Claude Code
openreadout mcp --install claude-desktop  # Claude Desktop
openreadout mcp --install codex           # OpenAI Codex
openreadout mcp --install cursor          # Cursor
openreadout mcp --install vscode          # VS Code / Copilot
openreadout mcp --install gemini          # Gemini CLI

Windsurf, Zed, Continue, and Cline are supported too. The server exposes 15 tools (openreadout_info, openreadout_check, openreadout_preview, openreadout_export, openreadout_analyze, ...) over JSON-RPC — no shell access needed.

Claude Code Plugin

Installs the MCP server and the skill together:

/plugin marketplace add openreadout/agent-plugins
/plugin install openreadout@openreadout

Gemini CLI Extension

gemini extensions install https://github.com/openreadout/agent-plugins

Codex Plugin

codex plugin marketplace add openreadout/agent-plugins
codex plugin add openreadout@openreadout

The Claude Code plugin, the Gemini CLI extension and the Codex plugin each add the skill and the MCP server. They come from the small openreadout/agent-plugins repository, which each release updates. The server runs the openreadout binary from your PATH, so install it first.

Agent Skill

openreadout self skill --install claude   # ~/.claude/skills/openreadout
openreadout self skill --install agents   # ~/.agents/skills/openreadout (Codex, Cursor, Copilot, Gemini CLI)

The skill source is in skills/openreadout.

Why your agent will thrive on OpenReadout

  • Deterministic JSON output — every command supports --json with published schemas. No regex parsing, no scraping stdout.

  • Fixed exit codes — 0 ok, 1 error, 2 usage, 3 unknown format, 4 corrupt file, 5 I/O, 6 unsupported feature. Agents branch on the code, not on the message.

  • Self-healing errors — every error carries a hint that says what to do next. Agents self-correct without human intervention.

  • Assurance on every answer — each result says whether files like it were validated against an independent reader. Agents know when to double-check.

  • Built-in preview renderer — preview writes a PNG the agent can look at. Agents can see the image, trace, or plate they are reasoning about.

  • Cheap metadata — info reads headers only, so a 100 GB file costs the same as a small one. --only returns just the fields asked for, saving tokens.

  • Safe by default — inputs are opened read-only and nothing connects to the network.

Error Recovery

# Agent asks for an image that does not exist
openreadout preview cells.lif --image 3 --json
{
  "ok": false,
  "error": {
    "code": "usage",
    "message": "usage error: image 3 not found (file has 1 images)",
    "hint": "Indices are zero-based; `openreadout info FILE --json` lists the images, traces (sweep_count, sample_count), tables (row_count) and spectra the file holds.",
    "exit_code": 2
  }
}

The agent follows the hint, lists the images, and picks the right index.

Python and R

Python — pip install openreadout returns metadata as dicts and pixels as NumPy, dask, or xarray arrays, with plugins for bioio and napari. Until wheels are published, run pip install . in a checkout. See the Python guide.

import openreadout

with openreadout.File("cells.lif") as f:
    f.images[0]["channels"]     # same keys as `info --json`
    stack = f.to_xarray(0)      # labelled with channel names and µm

R — the R package returns arrays and data frames. See the R guide.

Comparison

OpenReadout

Bio-Formats

bioio

czifile / nd2 / liffile

msconvert

Open source & free

✓ (MIT / Apache-2.0)

✓ (GPL)

✓ (plugins vary)

✓ (BSD)

✓ (vendor DLLs are not)

AI-native CLI + JSON + MCP

✓

✗

✗

✗

✗

Zero install (single binary)

✓

✗ (JVM)

✗ (Python)

✗ (Python)

✗

No vendor DLLs

✓

✓

✓

✓

✗

Integrity check

✓

✗

✗

✗

✗

Microscopy

✓

✓

✓

✓ (one format each)

✗

Mass spectrometry

✓

✗

✗

✗

✓

Ephys, flow, NMR, chromatography, plates, qPCR

✓

✗

✗

✗

✗

Cross-platform

✓

✓

✓

✓

Windows (or Wine)

Validation

Readers are tested against about 1,500 public instrument files. Each file's geometry, metadata, and plane hashes are compared with independent libraries (czifile, nd2, liffile, Bio-Formats, FlowIO, pyABF, and others), and pixel data must match exactly. See Validation.

Every reader was written from public files and permissively licensed documentation — no vendor SDKs, headers, DLLs, or GPL source code. See the clean-room policy.

Documentation

The documentation has guides for every command and format:

Privacy

OpenReadout runs on your computer, makes no network connections and sends no telemetry. See PRIVACY.md.

License

Licensed under either the Apache License 2.0 or the MIT license, at your option. OpenReadout is not affiliated with any instrument vendor; see TRADEMARKS.md.

Bug reports and contributions are welcome on GitHub Issues. See CONTRIBUTING.md, and read the clean-room policy before working on a reader.

Images and demo files come from the public test corpus (corpus/manifest.toml), used under their licences: mouse section, Zeiss sample images for Bio-Formats (Zenodo 10577621, CC-BY-4.0); Convallaria lambda scan, Maria Manuela Azevedo (Zenodo 14976703, CC-BY-4.0); BaTiO3 STEM, Rama Vasudevan and Gerd Duscher (Zenodo 8190744, CC-BY-4.0); H&E QPTIFF, PerkinElmer via the OME sample images (CC-BY-4.0); qPCR, the RDML R package (MIT); MS2, ProteoWizard test data (Apache-2.0); HPLC, cheminfo (MIT); EPR, EasySpin (MIT); patch clamp, pyABF (MIT); GC-FID, entab (MIT); terminal demo, Allen Institute for Cell Science (BSD-3-Clause). demo.tape regenerates the demo.


If you find OpenReadout useful, please give it a star on GitHub — it helps others discover the project.

Available Tools

15 tools
openreadout_analyzeA
Read-onlyIdempotent

An analysis with a documented method, picked by kind. options holds that kind's settings: the schema lists every kind's, and an option the kind does not take is an error naming the ones it does. openreadout_batch runs the same kinds and options over many files.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute or working-directory-relative path to the instrument file (kind=gate: the FCS file, or the gating file itself to describe it).
kindYesThe analysis. `peaks` = chromatographic peaks (detector trace, TIC, XIC, SRM): rt, area, height, widths, tailing, plates, resolution, S/N, area % (purity); compound lists; bands and regions of IR/Raman/UV-Vis/NMR spectra; `chromatogram` = TIC, BPC, XIC (mz + ppm), SRM transitions, stored or detector chromatograms: apex time and intensity, integral, thinned arrays ('when does m/z X elute'); `nmr-peaks` = NMR peak list (ppm, height, width, S/N) and region integrals, from the processed spectrum or the FID; `ephys-features` = patch-clamp: action potentials per sweep, rheobase, f–I slope, input resistance, tau, capacitance, sag; voltage-clamp holding current and access resistance; `spikes` = extracellular spike detection per channel: counts, rates, times; `qpcr` = qPCR (RDML, .eds, .rex): Cq and Tm per well × target, ΔΔCq fold changes, standard curves; `assay` = plate-reader assays: per-well values, standard curves with back-calculated concentrations, dose-response IC50/EC50, kinetics, growth, Z′; `gate` = flow-cytometry gating from a FlowJo workspace or Gating-ML file: population counts, percentages and medians.
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).
optionsNoThe options of `kind` (all optional; the schema lists every kind's).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior, so the bar is lower, and the description still adds non-obvious behavior: the `options` object is kind-discriminated and an option the kind does not take raises an error naming the valid ones. That is a genuine failure-mode disclosure not present in the annotations. It does not mention size limits, path-resolution expectations, or other runtime constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler; the purpose statement is front-loaded and the batch alternative is deferred appropriately. It is tightly written, though the opening sentence is abstract enough that it reads more as a framing statement than a definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the annotations carry the safety profile, so the description's remaining job is the kind/options contract, which it covers well along with the multi-file alternative. The only gap is that the resource and result nature are left entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline would be 3, but the description adds real meaning: it explains that `options` is a discriminated union where only the selected kind's settings are valid, which resolves the confusing `anyOf`-over-all-kinds shape in the schema. It does not clarify the `strict` or `file` semantics, which the schema already handles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says it produces 'an analysis with a documented method, picked by kind', which is a clear action but never names the resource being analyzed (instrument files) or what the analyses actually produce. The real specifics of purpose live in the `kind` enum, not the description, so an agent must open the schema to know what this tool is for. It does route to a sibling, which keeps it above pure vagueness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly names the alternative and the condition that selects it: 'openreadout_batch runs the same kinds and options over many files', implying this tool is for a single file. That is a real when-to-use-ex-batch signal. It stops short of an explicit 'use this when X' statement or any exclusions beyond the batch routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_batchA
DestructiveIdempotent

One measure over many files as one tidy table, optionally joined to sample sheets or plate maps (the join key is chosen from the data and reported in joins[] with unmatched rows) and summarized by group (by; test against a control). Analyses take the same options as openreadout_analyze. Inputs: files, directories, globs or an index query. Returns at most limit rows (page with offset; output writes them all to a file); a failing file is a row with error. measure=summarize regroups a table written by an earlier output.

ParametersJSON Schema
NameRequiredDescriptionDefault
byNoSummarize by these columns (e.g. `["condition"]`): n, mean, sd, sem, median, min, max, cv_percent per group; channels/parameters/populations stay apart automatically.
keysNoFixed join keys `SHEET_COLUMN=FIELD` (path, file, stem, sample_id, sample_name, barcode, well, position, run_order, column:NAME).
testNo`welch` or `mann-whitney` against `control` (a value of the first `by` column).
limitNoRows to return (default 20, max 500; 0 with `output`: the rows are in the file).
queryNo… and the query (search syntax, e.g. "format=fcs channel~CD4").
whereNoRow filters `COLUMN=VALUE` / `COLUMN!=VALUE` (e.g. "parameter=FITC-A").
fieldsNoColumns to keep (info: index field names, or `["all"]`).
inputsNoFiles, directories or glob patterns (summarize: the one table file).
offsetNoFirst row returned.
outputNoWrite the full table (or the summary, with `by`) to this .csv/.tsv/.jsonl/.json/.parquet file (verified; never overwrites without `overwrite`).
valuesNoValue columns to summarize (default: the measure's main values).
controlNoThe control group.
formatsNoOnly these format ids (e.g. `["fcs"]`).
measureYesWhat to measure per data set. `stats` = pixel statistics per image × channel (per=well for screening plates); `trace` = statistics per trace × sweep × channel; `table` = FCS: events, median, mean, sd, min, max per parameter; plate reads: one row per well; `info` = one row of header metadata per data set (fields); `spectra` = one row per MS scan header (options: the openreadout_spectra filters); `peaks` = openreadout_analyze kind peaks (options); `chromatogram` = openreadout_analyze kind chromatogram (options); `assay` = openreadout_analyze kind assay (options); `nmr-peaks` = openreadout_analyze kind nmr-peaks (options); `ephys-features` = openreadout_analyze kind ephys-features (options); `spikes` = openreadout_analyze kind spikes (options); `qpcr` = openreadout_analyze kind qpcr (options); `gate` = population counts, percentages and medians (workspace or gatingml); `summarize` = group statistics (by) of the one table file in inputs, written by an earlier output.
optionsNoThe measure's options. Analyses: the openreadout_analyze options of that kind, plus `rows` (which record list becomes rows: peaks peak|compound|chromatogram|band|region, nmr-peaks peak|integral|spectrum, ephys-features sweep|cell|spike, qpcr record|rq|standard_curve, assay wells|samples|compounds|kinetics|growth|quality). stats: image, select, level, per (channel|image|plane|well|field), wells, mip (z|t). trace: trace, sweep, channels. table: table, parameters, compensate, transform, workspace or gatingml, sample. gate: workspace or gatingml, sample, populations, medians, table. info: fields. spectra: the openreadout_spectra filters.
exact_byNoGroup by exactly `by` (do not add the measurement columns that vary).
overwriteNoReplace an existing output file.
recursiveNoWalk sub-directories.
replicateNoAverage rows within each replicate first.
worksheetNoWorksheet of an XLSX sheet.
from_indexNoAlso take the data sets an index query selects: the index directory …
sample_sheetsNoSample sheets (CSV/TSV/XLSX) or plate layouts (plate-map grids) to join; the key is chosen from the data (file name, path, sample id recorded in the file, well, barcode, vial, run order) and reported under joins[].

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=true; the description adds genuine behavioral context beyond them — a failing file becomes a row with error, results are capped by limit/offset, output writes the full table to a file, and joins are reported in joins[] with unmatched rows. This meaningfully lowers the need to inspect the schema for error/pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the body is a single dense run-on paragraph with nested parentheticals that is hard to parse. Most content earns its place for a 22-param tool, yet the structure could be broken into clearer segments.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 22-parameter, nested-object tool with an output schema, the description covers the overall workflow, join behavior, pagination/limit semantics, error-row handling, and the summarize mode. Output-schema presence excuses it from explaining return values, and the remaining gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 22 parameters thoroughly (including join keys, by, measure enum, and options). The description restates the input model (files, directories, globs, index query) and join-key selection, but adds little syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope: 'One measure over many files as one tidy table,' which cleanly distinguishes the batch tool from single-dataset siblings like openreadout_table, openreadout_stats, and openreadout_analyze. An agent can identify this as the many-files aggregator without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Establishes the batch context ('one measure over many files') and routes options to openreadout_analyze ('Analyses take the same options as openreadout_analyze'), plus notes measure=summarize regroups a prior output. It lacks an explicit when-not-to-use or a direct 'prefer X over Y' statement, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_checkA
Idempotent

Integrity check: ok plus findings (severity, code, message) for truncation, missing planes or parts, bad blocks; a plate folder reports missing files per well; a file still being written reports acquisition_in_progress. against compares two files instead (e.g. a source and its export): metadata differences, geometry, channels, per-plane hashes or max |difference| within tolerance. report=true builds a privacy-reviewed diagnostic bundle for the maintainers (no data values; free text only with include_text) with the issue_url to file it at.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute or working-directory-relative path to the file or data-set directory.
imageNoagainst: only this image index (in both files).
ignoreNoagainst: JSON pointers (into openreadout_info output) to leave out of the metadata diff; `*` matches one segment.
outputNoreport: also write the bundle to this new local file. Default: only return it.
reportNoBuild a privacy-reviewed diagnostic bundle for the maintainers instead (for a file that is refused, fails, or is not validated).
selectNoagainst: plane selection strings such as `c=0`, `z=2-5`, `t=0,3`. A second file holding only the selected planes (an export with the same selection) is matched to them in order.
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).
againstNoCompare with this second file instead (e.g. an OME-TIFF export of `file`): metadata differences, geometry, channels, physical sizes and per-plane hashes.
no_pixelsNoagainst: metadata and geometry only, no pixels.
toleranceNoagainst: largest absolute sample difference that still counts as equal. Default: bit-identical planes (same xxh3-128).
include_textNoreport: keep free text from the file (sample, image and channel names, comments); the path and personal data are still replaced. Default false: only with the user's consent.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations give readOnlyHint=false/destructiveHint=false/idempotent, and the description usefully explains the write behavior behind that (report writes a bundle to `output`, default only returns it) plus refusal semantics (strict → error with exit_code 6 and a hint). It also discloses the privacy model (no data values; free text only with include_text). Missing: any note on cost/runtime for large plates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, which is good, but the rest is a single run-on paragraph of semicolon-joined clauses mixing three modes. Every clause carries information, yet the mode boundaries are hard to parse on a first read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the description still covers all three operating modes, privacy handling, and the refusal path. Adequate for an 11-parameter tool; only the sibling-tool boundary is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it groups parameters by mode (plain check vs `against` vs `report`), explains what `against` compares (metadata, geometry, channels, per-plane hashes or max |difference| within tolerance), and clarifies `strict` and `include_text` behavior beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Integrity check: ok plus findings ... for truncation, missing planes or parts, bad blocks') and explicitly names the two alternate modes (against-comparison, report bundle). An agent can distinguish it from openreadout_info/export/analyze without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'instead' framing tells the agent when to switch modes: pass `against` to compare two files, pass `report=true` for a refused/failed file. It also notes acquisition_in_progress for in-progress writes. It does not, however, explicitly contrast with sibling tools like openreadout_info or openreadout_export, which the comparison mode overlaps with.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_exportA
DestructiveIdempotent

Convert to an open format: a new file, read back and verified; the source is never touched. format: ome-tiff (images, default; pyramidal when the source is), ome-zarr (multiscale; plates as OME-NGFF HCS), parquet or arrow (tables, traces with every sweep, spectra=true for MS points), mzml, nwb (electrophysiology), jcamp (NMR, IR/Raman, chromatograms), asm (plate readers), rdml (qPCR). attachment writes one embedded attachment (thumbnail, label, slide preview) as stored. Sends progress with a progressToken.

ParametersJSON Schema
NameRequiredDescriptionDefault
runNomzML, Parquet, Arrow: spectra run index (default 0).
fileYesAbsolute or working-directory-relative path to the instrument file.
rowsNoParquet, Arrow, NWB, JCAMP-DX: rows (samples of each sweep for traces) `A-B`, `A-` or `A`, zero-based and inclusive.
imageNoImages: only this image index.
levelNoImages: export this source pyramid level as the full resolution (default 0).
sweepNoParquet, Arrow, NWB, JCAMP-DX: only this sweep (default: every sweep).
tableNoParquet, Arrow: this table index (FCS data set, plate read, event or peak table).
traceNoParquet, Arrow, NWB, JCAMP-DX: this trace index (NWB default: every trace).
wellsNoOME-Zarr of a multi-well plate: only the fields of these wells (`C05`).
formatNo`ome-tiff` (default for images; one BigTIFF file), `ome-zarr` (OME-NGFF 0.5 / Zarr v3 directory with a pyramid), `mzml` (indexed mzML 1.1.0, for mass-spectrometry files; the default for them), `asm` (Allotrope Simple Model plate-reader JSON; plate-reader exports only), `rdml` (RDML 1.3; qPCR files: RDML, Applied Biosystems .eds, Rotor-Gene .rex), `parquet` or `arrow` (Arrow IPC file) for a table, a trace (all sweeps) or mass spectra (`spectra=true`), `nwb` (NWB 2.x; electrophysiology traces) or `jcamp` (JCAMP-DX 5.01; NMR FIDs and spectra, other 1-D spectra and chromatograms).
outputNoOutput path; defaults to the input with `.ome.tiff`, `.ome.zarr`, `.mzML`, `.asm.json`, `.rdml`, `.parquet`, `.arrow`, `.nwb` or `.jdx`.
regionNoImages: export only this rectangle `{x, y, width, height}` (pixels of `level`) of every plane, read tile by tile (cut a field out of a whole-slide image).
selectNoImages: plane selection strings such as `c=0`, `z=2-5`, `t=0,3`.
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).
spectraNoParquet, Arrow: export the mass spectra of `run` (one row per point, plus a `<name>.scans.parquet` per-scan summary) instead of a table or trace.
centroidNomzML, Parquet, Arrow: write the instrument's centroid lists instead of profiles where a scan has both.
overwriteNoReplace an existing output file or directory.
attachmentNoWrite this embedded attachment instead (a CZI's `Thumbnail`, `Label` or `SlidePreview` image, `TimeStamps`; names from openreadout_info view=structure, kind attachment), or `#<index>`.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and readOnlyHint=false, and the description reconciles this correctly: 'the source is never touched' while overwrite is parameter-controlled, and the output is 'read back and verified'. It also discloses progressToken progress reporting, adding real context beyond the annotations. It stops short of describing error/refusal behavior beyond what the schema's `strict` note already covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core guarantee (new file, source untouched) is front-loaded, then formats, then attachment, then progress. Dense but every clause carries information for an 18-parameter tool. Slightly list-heavy, but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 18-parameter, no-output-schema export tool, the description covers the main modes (format conversion, attachment extraction, progress) and the safety guarantee. It doesn't need to explain return values since none are declared, though it could say more about failure/verification outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description restates the format menu and adds a couple of nuances not in the schema (pyramidal output mirroring the source, plates as OME-NGFF HCS), but most parameter meaning is already carried by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'Convert to an open format: a new file, read back and verified; the source is never touched.' This clearly frames it as a file-export tool and the format enumeration reinforces scope. It does not explicitly contrast itself with siblings like openreadout_preview or openreadout_analyze, so a 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (when you need a converted, verified copy on disk, or a stored attachment via `attachment`), but there is no explicit 'use this instead of X' routing against the many siblings. Usage is implied by the format/attachment modes rather than stated as guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_formatsB
Read-onlyIdempotent

Supported formats with read/write support, confidence and known gaps (what is not decoded).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
formatsYesEvery registered format, in detection order.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds genuine content about what the listing reports (per-format read/write capability, confidence, undecoded gaps), but adds nothing about freshness or scope of the format registry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the parenthetical defining 'known gaps' is the only elaboration. It is dense but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value shape need not be explained, and the tool is parameterless and read-only. The description covers the content of the payload adequately; only the sibling routing question remains open.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. No parameter-level confusion is possible.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (supported formats) and enumerates what the listing contains: read/write support, confidence, and known gaps. It is far more concrete than a tautology, but it does not distinguish itself from the similarly named sibling openreadout_info, leaving the agent to guess which discovery tool to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus alternatives, nor any prerequisite or trigger condition. With an ambiguous sibling like openreadout_info in the toolset, the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_indexA
Idempotent

Catalog every data set under roots into index_dir (Parquet: one row per data set with format, size, sample, instrument, technique, method, operator, start, dimensions, channels, objective, events, scans, integrity status, personal-data flags). Headers only and resumable: a call stops after max_files/max_seconds with complete=false, and the same call again continues. openreadout_search queries the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
piiNoLook for personal data (default true).
checkNoIntegrity check per data set: `headers` (default; structure only), `full` or `none`.
rootsYesDirectories (or files) to crawl.
excludeNoGlob patterns of names or root-relative paths to skip.
restartNoDiscard an interrupted crawl instead of resuming it.
threadsNoWorker threads (default: the number of CPUs).
index_dirYesIndex directory to create or update (never inside a root).
max_filesNoStop after this many files in this call (default 20000); call again to continue.
full_rescanNoRead every file again instead of reusing unchanged records.
max_secondsNoStop after this many seconds in this call (default 45); call again to continue.

Output Schema

ParametersJSON Schema
NameRequiredDescription
piiYesPersonal-data totals.
nextNoWhat to do next, when the crawl is not complete.
toolYesThe tool and version that wrote it.
bytesYesBytes of every file seen.
crawlYesCrawl counters, cumulative over resumed sessions.
filesYesRows of `files.parquet`.
rootsYesThe directories (or files) crawled, absolute.
yearsYesData sets per acquisition year (`unknown` when the file records no date).
tablesYesThe tables.
changesYesData sets per change (`new`, `changed`, `unchanged`, `moved`) and `removed`.
formatsYesData sets and bytes per format id.
completeYesTrue when the crawl walked every root to the end. False after a cap (`--max-files`, `--max-seconds`) or an interruption: the tables then hold what was read so far plus the previous run's records for the rest, and rerunning the same command continues.
datasetsYesData sets (rows of `experiments.parquet`) after multi-file grouping.
familiesYesData sets and bytes per family.
problemsYesProblems per category (`integrity`, `readability`, `walk`, `pii`).
settingsYesOptions of the crawl.
index_dirYesAbsolute path of the index directory.
started_atYesWhen this run started (ISO-8601 UTC).
updated_atYesWhen the tables were written.
files_per_sYesItems per second over the crawl's wall time.
check_statusYesData sets per check status.
schema_versionYes[`INDEX_SCHEMA_VERSION`].
unknown_extensionsYesFiles no reader recognises, per lower-case extension (`(none)` for no extension).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and the description adds genuinely useful behavior: headers-only pass, resumable with complete=false, restart/full_rescan semantics, and PII discovery. This goes beyond the annotation flags, though it does not describe write location constraints beyond 'never inside a root' (which is in the schema) or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and output shape, then the resumption contract, then the sibling handoff. Dense but every clause carries information; the long parenthetical field list is the only slightly heavy element, and it is arguably warranted given no output schema is described in prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the operation, resumption model, and downstream tool. For a 10-parameter mutation tool it is nearly complete, missing only explicit guidance on choosing between resuming and restart/full_rescan and any notion of cost or duration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description earns above that by explaining the cross-parameter interaction semantics of max_files/max_seconds ('a call stops ... and the same call again continues'), which the schema fields describe only individually.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Catalog every data set under roots into index_dir') and immediately characterizes the artifact produced (Parquet, one row per data set with listed fields). It also names the sibling that consumes the result (openreadout_search), so an agent can place it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: build an index first, then query it via openreadout_search, and explains the resumable calling pattern (stop at max_files/max_seconds, call again to continue). It does not explicitly state when not to use it (e.g., vs. openreadout_watch or openreadout_check on an already-indexed share), but the intended flow is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_infoA
Read-onlyIdempotent

What an instrument file or data-set directory holds, from its headers (fast on huge files): format, images[] (sizes, pixel type, µm/px, channels and dyes, objective), tables[], traces[], spectra[] (MS runs), plate, experiment (sample, instrument, method, operator, start) and assurance. view=summary attaches a ~384 px thumbnail of image 0 with full-resolution pixel rulers. ask answers a question in words, naming the fields it used.

ParametersJSON Schema
NameRequiredDescriptionDefault
askNoA question in plain words ("what was the gradient?", "which channel is DAPI?", "how many MS/MS scans?"): answered under answers[] with the fields each came from. Implies view=explain.
fileYesAbsolute or working-directory-relative path to the instrument file or data-set directory.
viewNoWhat to return. `summary` = the header summary (default); `full` = the summary with per-frame records (time stamps, stage positions, exposure) and per-field provenance; vendor=true adds the vendor's raw metadata tree; `structure` = the container: images, blocks/segments/chunks, attachments (label, thumbnail, previews), pyramid levels, with offsets and sizes; `explain` = a plain-English account with caveats and suggested next steps (ask answers questions); `format` = the format only, from the file's signature (cheapest; works on damaged files).
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate (assurance.strict_refuses). Default: the server's setting (OPENREADOUT_STRICT; off).
vendorNoview=full: include the vendor's raw metadata tree (can be megabytes). Default false.
thumbnailNoview=summary: attach a small picture of image 0 (~384 px, with full-resolution pixel rulers) as image content. Default true; skipped (with a note) when it would decode too much.
max_framesNoview=full: per-frame records per image under images[].extra.frames. Default 100; 0 omits them; -1 embeds all.
max_imagesNoList at most this many images (0 = all). Default: all, except screening plates (16; their field images are all alike and plate.wells[].images indexes every field).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed world), so the bar is lower, and the description still adds real behavior: header-only reads are fast even on huge files, thumbnails are skipped with a note when decoding would be too costly, vendor trees can be megabytes, and strict mode refuses unvalidated values. Return-shape details are largely deferred to the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the resource and the cost characteristic ('fast on huge files'), then the payload inventory, then the view=summary and ask modifiers. Slightly packed but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with a full output schema, the description supplies enough orientation: what the header report contains, the cost profile, and the two notable modifiers. It need not explain return values, though it could say more about when to reach for it versus sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the per-parameter docs are already detailed, so baseline 3 applies. The description restates view=summary thumbnail behavior and that ask implies view=explain, which mirrors rather than extends the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource (an instrument file or data-set directory) and enumerates exactly what it reports: format, images, tables, traces, spectra, plate, experiment, assurance. It distinguishes itself from generic siblings like openreadout_stats or openreadout_preview by scoping to header-derived metadata, but it never names an alternative tool explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: 'fast on huge files' hints that this is the cheap first-look tool, and the view/ask semantics suggest follow-up queries. There is no explicit when-to-use or when-to-prefer-a-sibling guidance against the 15 other openreadout_* tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_previewA
Read-onlyIdempotent

A picture of the data as image content plus what was drawn: an image plane (default channel 0, middle z), a composite, a max projection (mip=z), a region (only the tiles needed are read), a trace or spectrum plot, or a plate heat map. Image pictures carry rulers in full-resolution pixels and a µm scale bar, so a region read off the rulers zooms in.

ParametersJSON Schema
NameRequiredDescriptionDefault
lutNo`gray` or `channel-color` (default: gray for one channel, channel colours for composites).
mipNoMaximum-intensity projection axis: `z` (or `t`), over the selected range if any.
runNoSpectrum preview: run index.
axesNoImage previews: coordinate rulers in full-resolution pixels and a scale bar around the picture (default true); false = the bare plane.
fileYesAbsolute or working-directory-relative path to the instrument file.
gridNoImage previews: faint grid lines at the ruler ticks over the data (default false).
scanNoSpectrum preview: instrument scan number instead of an index.
imageNoImage index (default 0).
levelNoPyramid level to read (default: the level nearest max_size, or the level at which `region` renders at about max_size).
sweepNoTrace preview: sweep index.
tableNoPlate preview: table index.
traceNoTrace preview: trace index (see openreadout_info traces[]).
columnNoPlate preview (long layout): value column name.
formatNo`png` (default) or `jpeg`.
regionNoZoom: only this rectangle, `{x, y, width, height}` in full-resolution pixels (or in the pixels of `level` when a level is given). Whole-slide images: look at the overview first, then at regions of it.
selectNoPlane selection such as `c=1`, `z=4`, `t=0` or combined `c=0,1,z=2`. Default: c=0, the middle z, t=0. Several channels imply a composite.
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).
centroidNoSpectrum preview: the instrument's centroid list instead of the profile.
channelsNoTrace preview: channel indices (default: the first 8).
contrastNo`auto` (default), `min-max`, `percentile:LO,HI` (e.g. `percentile:1,99`) or `raw`.
max_sizeNoLongest side in pixels, rulers included (default 768, max 2048).
spectrumNoSpectrum preview: zero-based spectrum index.
compositeNoBlend all (or the selected) channels additively in their colours.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hintNoWhat to do next with the picture (look at it; zoom with a region read off the rulers).
kindYes`image`, `trace`, `spectrum` or `plate`.
pathYesThe input file.
xxh3Yesxxh3-128 of the encoded bytes, 32 hex chars (identical across runs for the same request).
bytesYesSize of the encoded image.
imageNoPresent for image previews.
notesNoAnything the caller should know (fallbacks, caps, implied options).
plateNoPresent for plate previews.
traceNoPresent for trace previews.
widthYesImage width in pixels.
formatYesFormat id of the input file.
heightYesImage height in pixels.
outputNoThe written file (CLI); absent when the preview is returned inline (MCP image content).
encodingYes`png` or `jpeg`.
spectrumNoPresent for spectrum previews.
verifiedYesTrue once a written file was read back and matched.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description still adds real behavioral context: image pictures carry rulers in full-resolution pixels plus a µm scale bar, and a region read only touches the tiles needed, which tells the agent the operation is cheap and how coordinates relate to what is drawn.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is effectively two sentences: an enumerated list of preview kinds followed by a note on rulers and zoom. The purpose is front-loaded and nothing is obviously padded, though the first sentence is a dense run-on enumeration rather than crisp structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and full parameter documentation, the description need not explain return values, and it does orient the agent across the tool's several preview modes. It stops short of tying each mode to its required parameter cluster (trace/sweep/spectrum/table) or explaining the strict/contrast/lut behaviors, but the schema fills those gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 23 parameters and the baseline is 3. The description reinforces a few pieces (default channel 0, middle z, mip=z, region in full-resolution pixels) but adds little syntax or interaction detail beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb and resource: it renders a picture of the data (image plane, composite, mip, region, trace/spectrum plot, plate heat map). That gives an agent a clear sense of what openreadout_preview produces, though it never explicitly contrasts itself with visual-adjacent siblings like openreadout_export or openreadout_trace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'a region read off the rulers zooms in' hints at an iterative look-then-zoom workflow, and the modes named map implicitly to parameter sets. There is no explicit when-to-use, when-not-to-use, or pointer to an alternative tool (e.g. export vs preview), so an agent must infer routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_spectraA
Read-onlyIdempotent

Mass spectra of an MS run. Without scan, index or nth: the scan headers without decoding peaks (scan number, MS level, retention time, polarity, precursor m/z and charge, isolation window, activation, collision energy, filter, stored TIC), filtered and paged; every match is counted (matched, ms_level_counts). With scan, index, or ms_level + nth: that one spectrum's mz[] and intensity[] with its metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
nthNoOne spectrum: with ms_level, the nth spectrum of that level, counting from 1 (ms_level=2, nth=1 = the first MS/MS scan; MS1 and MS/MS scans interleave).
ppmNoPrecursor tolerance in ppm instead of 0.01 m/z.
runNoRun index for files with several runs (Sciex samples; default 0).
fileYesAbsolute or working-directory-relative path to the mass-spectrometry file.
scanNoOne spectrum: the scan number as the instrument counts it (1-based in Thermo files; mzML `scan=N` native ids, mzXML `num`, timsTOF spectrum position + 1).
indexNoOne spectrum: the zero-based spectrum index.
limitNoList at most this many matching scans (default 100, at most 5000; 0 only counts); all are counted.
chargeNoOnly precursors of this charge state.
offsetNoSkip this many matching scans (paging; default 0).
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).
centroidNoOne spectrum: the stored centroid list instead of the profile when a scan has both.
ms_levelNoOnly scans of this MS level (1 = full scans, 2 = MS/MS); with nth, the level to count in.
polarityNoOnly `positive` or `negative` scans.
rt_rangeNoRetention-time window `[start, end]` in minutes.
activationNoOnly this activation: HCD, CID, ETD, ... (case-insensitive).
max_pointsNoOne spectrum: at most this many points (default 2000; point_count reports the full size).
scan_filterNoOnly scans whose filter string contains this text (case-insensitive).
precursor_mzNoOnly MS/MS scans whose precursor m/z is within 0.01 (or `ppm`) of this.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful behavior: the list mode returns headers 'without decoding peaks' (a cost/perf hint), that every match is counted even when paged, and that the detail mode returns mz[]/intensity[] with metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the resource then the mode split, with no filler. Slightly long clauses but every element (header fields, paging, counts vs mz[]/intensity[]) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 18-parameter tool with an output schema present, the description covers the critical behavioral fork (list vs single spectrum) and the counting/paging semantics. Return-value detail is delegated to the output schema, which is appropriate, leaving only minor gaps around error/paging edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the per-parameter descriptions are already detailed (nth semantics, ppm vs 0.01 m/z, limit cap of 5000, strict exit_code 6). The description adds little parameter-level meaning beyond restating which params select the detail mode, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (mass spectra of an MS run) and defines two distinct retrieval modes by parameter combination, so an agent knows exactly what this tool yields. It does not explicitly contrast itself with siblings like openreadout_preview or openreadout_search, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the trigger conditions clearly: omitting scan/index/nth gives paged scan headers, while supplying scan, index, or ms_level+nth returns a single decoded spectrum. That is explicit conditional routing, though it offers no guidance on when to prefer this tool over sibling tools such as openreadout_preview.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_statsA
Read-onlyIdempotent

Pixel statistics per channel and image: count, min, max, mean, std, percentiles p1–p99, zero and saturated fractions (at the detector's 2^bits−1 when recorded), histogram; RGB images also per component. per=plane adds every plane; per=well or field gives rows per well of a screening plate. mip='z' measures the maximum-intensity projection; whole slides: a coarser level and/or a region.

ParametersJSON Schema
NameRequiredDescriptionDefault
mipNoMaximum-intensity projection first: `z` (per image, channel and time point: the largest value of each pixel over the selected z planes) or `t`; the statistics are of the projections. The axis of a maximum-intensity projection (`stats --mip z`). `z` = Along z: one projected plane per image, channel and time point; `t` = Along t: one projected plane per image, channel and z.
perNoRows. `channel` = channel and image aggregates (default); `plane` = also one entry per plane (image, c, z, t); `well` = a screening plate (pass the plate folder or index, never one TIFF): one row per well × channel over the well's fields; `field` = a screening plate: one row per well × field × channel.
binsNoHistogram bins (0 = none). Default 32, at most 65536.
fileYesAbsolute or working-directory-relative path to the instrument file, or a plate's index file or folder (Harmony `Index.idx.xml`, ImageXpress `.HTD`, CellVoyager `MeasurementData.mlf`, OME-Zarr plate).
imageNoOnly this image index.
levelNoPyramid level: 0 = full resolution (default).
scaleNoHistogram spacing: `linear` (default) or `log`. Spacing of histogram bin edges. `linear` = Equal-width bins from `min` to `max` (integer data: to `max + 1`, so each bin holds whole values); `log` = Geometrically spaced bins from the smallest positive value to `max`; values ≤ 0 are counted in `nonpositive`. Suited to data spanning decades (photon counts, spectra).linear
wellsNoper=well|field: only these wells (`C05`, `c5`).
regionNoOnly this rectangle of each plane, `{x, y, width, height}` in the pixel coordinates of `level` (e.g. the 512 x 512 centre of a whole-slide image).
selectNoPlane selection strings such as `c=0`, `z=2-5`, `t=0,3`. The aggregates cover exactly the selection.
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the read-only, idempotent, non-destructive profile, so the bar is lower. The description still adds genuine behavioral context beyond that: saturated fractions are computed at the detector's 2^bits−1 only 'when recorded', and it discloses what the aggregates cover. No permissions or performance limits are mentioned, but the conditional-recording caveat is real added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core output metrics and then branches into the per/mip/whole-slide cases. It is dense and telegraphic with heavy semicolon use, but every clause carries information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 11 parameters, a rich output schema, and full annotations, the description covers the main operational branches (per modes, mip, region/level for slides) adequately. Return-value format is delegated to the output schema, which is appropriate, and nothing critical for correct invocation appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters thoroughly. The description largely restates the per and mip semantics that the schema already spells out, adding little parameter meaning beyond the baseline. A 3 is appropriate when the schema carries the semantic load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation (compute pixel statistics) on a specific resource (per channel and image) and enumerates the metrics returned (count, min, max, mean, std, percentiles, zero/saturated fractions, histogram). It is clearly distinguishable from sibling tools like openreadout_spectra or openreadout_table, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives conditional usage for options (per=plane adds every plane; per=well/field for screening plates; mip='z' for projections; coarser level/region for whole slides), which is helpful routing within the tool. However, it offers no guidance on when to choose this tool over alternatives such as openreadout_spectra or openreadout_table, leaving cross-tool selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_tableA
Read-onlyIdempotent

Rows of a table {columns, labels, rows, total_rows, truncated}: FCS events ($PnN columns), plate reads (well, row, col, read, wavelength_nm, time_s, value), event, spike or peak tables (including the vendor's own integration results). Values are raw; FCS can be compensated, transformed and given 0/1 population columns from a workspace or Gating-ML file. filter with count=true counts the events meeting conditions.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute or working-directory-relative path to the instrument file.
countNoReturn only the count (`filter.matched_rows`, `total_rows`, `percent`), no rows.
tableNoTable index (FCS data set or plate read; see openreadout_info → tables[]). Default 0.
filterNoRow conditions `COLUMN OP NUMBER` (OP: > >= < <= == !=; COLUMN: $PnN, $PnS or `gate:<path>`), all of which must hold, e.g. `["FITC-A > 1000"]`. Tested on the values returned (after compensate/transform when given) over the whole table; `filter.matched_rows` counts every match and first_row/max_rows page through them.
sampleNoWorkspace sample name or id (default: matched by file name or $FIL).
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).
gatingmlNoFCS only: Gating-ML 2.0 file (same uses as workspace).
max_rowsNoMaximum rows to return. Default 100 (fewer for wide tables: about 5000 values), capped at 10000.
first_rowNoZero-based first row. Default 0.
transformNoFCS only: transform after compensation: `logicle`, `arcsinh`, `hyperlog`, `log`, `linear`, `biex`, `flowjo-log`, `arcsinh-cofactor`, with parameters as `NAME:K=V,…` (`logicle:T=262144,W=0.5,M=4.5,A=0`, `arcsinh-cofactor:5`), or `workspace` for the FlowJo sample's per-parameter transforms.
workspaceNoFCS only: FlowJo workspace (.wsp) for compensate=gating, transform=workspace and populations.
compensateNoFCS only: compensate the values: `auto` (the gating file's matrix when one is given, else the file's $SPILLOVER/$SPILL/SPILL), `fcs` or `gating`. Values become FCS scale values first ($PnE, $PnG, $TIMESTEP).
populationsNoFCS only: populations (paths like `/Lymphocytes/Singlets` or unique names) whose 0/1 membership is appended as `gate:<path>` columns.
transform_parametersNoFCS only: parameters ($PnN) to transform; default every fluorescence parameter.

Output Schema

ParametersJSON Schema
NameRequiredDescription
pathYesThe input file.
rowsYesRow-major values; `rows[r][c]`. Non-finite values serialize as `null`.
tableYesTable index (see `info` → `tables[]`).
filterNoPresent with `--where`/`--count`: the row conditions and how many rows of the whole table meet them. `rows`, `first_row` and `truncated` then page through the matching rows.
formatYesFormat id of the input file.
labelsYesColumn labels (FCS `$PnS`), `null` where absent.
columnsYesColumn names (FCS `$PnN`).
first_rowYesZero-based index of the first row returned.
truncatedYesTrue when more rows follow the returned slice.
processingNoPresent when the values were processed (FCS scale values, compensation, transforms, population membership columns); absent for raw values.
total_rowsYesRows in the whole table.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/closed-world, so the safety profile is covered; the description adds genuine behavioral context beyond that: values are returned raw, compensation and transforms are optional and workspace/Gating-ML driven, and counts can substitute for rows. It stops short of noting truncation handling or paging limits in prose, which the schema covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Return shape is front-loaded, followed by supported table types and then the FCS processing and filter-count behavior. Dense but efficient, with parenthetical enumerations rather than filler; the mixed phrasing ('event, spike or peak tables') is slightly loose but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, yet the description still names the envelope fields and the table categories. For a 14-parameter, FCS-heavy tool the coverage of raw-vs-processed values and the count path is adequate, though it could say more about paging/truncation behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% across 14 parameters, so the schema carries the semantics and 3 is the baseline. The description does add conceptual linkage between compensate/transform/workspace/gatingml and mentions the count=true counting behavior, but most parameter-level detail is already in the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete resource and return shape ('Rows of a table {columns, labels, rows, total_rows, truncated}') and enumerates the table types it serves (FCS events, plate reads, event/spike/peak tables), so an agent knows what it gets back. It does not explicitly differentiate itself from siblings like openreadout_preview or openreadout_export, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the description notes that filtering plus count=true yields event counts, which hints at a counting use case, and it explains that FCS data can be compensated/transformed/population-annotated. There is no explicit when-to-use-this-vs-alternatives guidance, e.g. versus preview or export.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_traceB
Read-onlyIdempotent

One window of one sweep of a sampled signal or 1-D spectrum in physical units: electrophysiology sweeps, chromatography detector traces, ÄKTA curves, ITC thermograms, SPR sensorgrams, NMR FIDs and spectra (process=true turns an FID into a spectrum), IR/Raman/UV-Vis/CD spectra (each spectrum of a map is a sweep), qPCR curves, EPR, XRD, electrochemistry, thermal analysis. Per channel: min, max, mean, std and argmax_axis_value (the maximum's position on the axis: retention time, ppm, cm⁻¹, °2θ) over the window, plus the first max_samples values.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute or working-directory-relative path to the instrument file.
countNoWindow length in samples (statistics cover the whole window). Default: to the end of the sweep.
sweepNoSweep (episode or segment) index. Default 0.
traceNoTrace index (see openreadout_info → traces[]). Default 0.
strictNotrue: refuse (an error with exit_code 6 and a hint) values this file's assurance does not validate. Default: the server's setting (OPENREADOUT_STRICT; off).
processNoNMR: return FID traces as spectra processed by OpenReadout (group delay, apodization, zero filling, FT, stored or automatic phase, baseline, ppm axis).
x_rangeNoWindow on the trace's own axis instead of first_sample/count, `[a, b]` either order: cm⁻¹, nm, ppm, a chromatogram's retention time in its axis unit; seconds for signals without an axis.
channelsNoChannel indices to return. Default: all channels.
max_samplesNoSamples returned per channel. Default 200, capped at 10000.
first_sampleNoFirst sample of the window, zero-based within the sweep. Default 0.

Output Schema

ParametersJSON Schema
NameRequiredDescription
axisNoThe trace's abscissa (`info` → `traces[].extra.axis`: `{quantity, unit, first, step, …}`), e.g. the ppm axis of an NMR spectrum or a chromatogram's retention time in minutes; sample `i` of the sweep is at `first + i * step`. For a time axis this agrees with `start_s` and `sample_rate_hz` (`first` = `start_s` at `first_sample` 0, `step` = 1 / `sample_rate_hz`, in `unit`). Irregular axes (`irregular: true`) name the channel holding each sample's abscissa instead.
pathYesThe input file.
sweepYesSweep index.
traceYesTrace index (see `info` → `traces[]`).
formatYesFormat id of the input file.
start_sYesTime of the first window sample in seconds. Single-sweep traces use the trace's own clock: `info.traces[].start_s + first_sample / sample_rate_hz` (a chromatogram's retention time, negative when acquisition began before injection; an electrophysiology recording's clock). Multi-sweep traces are timed from the start of the sweep (`first_sample / sample_rate_hz`; `sweep_start_s` places the sweep in the recording). Irregularly sampled traces (`sample_rate_hz` 0 with a `time` channel) report that channel's first window value; traces without a time base (spectra) report 0.
x_rangeNoThe axis window asked for (`x_range`, low to high): axis units, or seconds for signals without an axis. `first_sample`/`sample_count` are the samples inside it.
channelsYesOne entry per requested channel.
truncatedYesTrue when fewer samples are returned than the window holds.
sweep_countYesSweeps in the trace.
first_sampleYesFirst sample of the window (zero-based, within the sweep).
sample_countYesSamples in the window (statistics cover all of them).
sweep_start_sNoTime of the sweep's first sample on the recording clock, seconds, when the file records it (`info.traces[].start_s` for a single sweep, `extra.sweep_starts_s[sweep]` for multi-sweep traces).
sample_rate_hzYesSamples per second.
sweep_sample_countYesSamples in this sweep.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description adds the windowing behavior and the process=true transformation semantics, but does not disclose anything about pagination limits, error behavior, or the assurance/strict interaction beyond what the schema already says.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core statement is front-loaded, which is good, but the first sentence is an extremely long enumerated list of instrument domains and the per-channel return explanation is a separate dense clause. It is informative but bloated; several domain examples are redundant for an agent that already knows this reads generic instrument traces.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, yet it spends most of its length on them anyway. For a 10-parameter tool it is reasonably complete on domain applicability but leaves the agent without routing guidance among the many sibling readers, so it is adequate rather than fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reinforces semantics the schema states (process turns an FID into a spectrum; the axis units for argmax_axis_value and x_range) and clarifies that 'each spectrum of a map is a sweep', but adds little that the parameter descriptions do not already contain.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource precisely ('one window of one sweep of a sampled signal or 1-D spectrum in physical units') and the extensive domain list (electrophysiology, chromatography, NMR, ITC, SPR, etc.) tells an agent exactly what kinds of files it operates on. It is clear about what is read, though it never explicitly contrasts itself with siblings like openreadout_spectra or openreadout_stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the domain enumeration and the process=true note give context for when the tool applies, but there is no explicit when-to-use vs when-not, and no routing to alternatives such as openreadout_stats or openreadout_table for the same signal data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openreadout_watchA
Read-only

Follow directories an instrument writes to: each call looks once and returns events since cursor (dataset_new, plane_new, scan_new, frame_new, dataset_complete, dataset_stalled, qc with qc=true: saturation, focus drift, dropped frames, TIC drop, error). Poll with the returned cursor; files are never locked.

ParametersJSON Schema
NameRequiredDescriptionDefault
qcNoFirst call only: evaluate the built-in live QC rules on each new plane or scan.
dirsYesDirectories (or files) to watch. The same list (in any order) continues the same watcher across calls.
limitNoMost events returned (default 200, max 5000); call again with the new cursor for more.
sinceNoFirst call only: report data sets modified since then (`10m`, `2h`, `1d`, an ISO-8601 time, or `all`). Default: the live window (5 minutes) back from now.
cursorNoReturn events with a sequence number greater than this (the `cursor` of the previous call); 0 or absent on the first call.
stall_after_sNoFirst call only: seconds without growth before `dataset_stalled` (default 120).

Output Schema

ParametersJSON Schema
NameRequiredDescription
moreYesTrue when more events are waiting (raise `limit` or call again).
cursorYesPass back as `cursor` to get only newer events.
eventsYesEvents after the given cursor, oldest first (the same objects `openreadout watch` prints under `data`).
droppedYesTrue when events after the cursor were dropped (the server keeps the last 10000).
trackedYesData sets the watcher remembers.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, idempotentHint=false), and the description usefully adds that each call is a one-shot look rather than a live stream, that polling uses the returned cursor, and that files are never locked. That is meaningful behavioral context beyond the annotations, though it omits rate/backpressure details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler, front-loading the core purpose and then the polling/event semantics. The event enumeration is long but informative and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description covers the essential polling model and non-locking guarantee. Combined with the rich schema, an agent has enough to call it correctly, though it could mention how to stop watching or cursor reset behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter (qc, dirs, limit, since, cursor, stall_after_s) in detail. The description only restates the cursor-polling pattern and the qc=true behavior, which the schema already provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (follow/watch) and resource (directories an instrument writes to) and enumerates the event types emitted, so an agent knows exactly what this tool does. It doesn't explicitly contrast itself with siblings like openreadout_index or openreadout_search, but the streaming/polling nature is distinctive enough that confusion is unlikely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational guidance for continued use ("Poll with the returned cursor") and a safety note ("files are never locked"), but says nothing about when to prefer this tool over alternatives such as openreadout_index or openreadout_check, and no explicit when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 15 tool updates
    • First observedopenreadout_analyze
    • First observedopenreadout_batch
    • First observedopenreadout_check
    • First observedopenreadout_export
    • First observedopenreadout_formats
    • First observedopenreadout_index
    • First observedopenreadout_info
    • First observedopenreadout_link
    • First observedopenreadout_preview
    • First observedopenreadout_search
    • First observedopenreadout_spectra
    • First observedopenreadout_stats
    • First observedopenreadout_table
    • First observedopenreadout_trace
    • First observedopenreadout_watch

TDQS

A3.9/5.0

Scored across 15 tools

Disambiguation4/5

Most tools are cleanly separated by data type or action (spectra for MS, trace for signals, table for tabular, stats for pixels, preview for pictures, check for integrity, export for conversion). However, openreadout_analyze is vaguely defined ('an analysis with a documented method, picked by kind') and overlaps conceptually with stats/table/trace/spectra, and info/formats/check have adjacent purposes that could occasionally be confused.

Naming Consistency5/5

Every tool uses the same lowercase snake_case convention with a uniform openreadout_ prefix, and each name is a single clear action or noun (analyze, check, export, index, search, spectra, stats). No mixed casing or verb-style drift.

Tool Count5/5

15 tools is well within the ideal range and each covers a distinct capability (analysis, batch, integrity, export, formats, indexing, info, linking, preview, search, spectra, stats, table, trace, watch). No redundant filler.

Completeness4/5

The surface is broad: read/inspect (info), validate (check), convert (export), catalog and query (index/search), link across instruments, visualize (preview), and analyze (stats/table/trace/spectra/analyze/batch), plus streaming (watch). Minor gaps around explicit write/edit or richer QC workflows, but the read-and-convert lifecycle is well covered.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Searches and fetches research datasets across Zenodo, DataCite (Dryad/Figshare/Dataverse/OSF), NCBI omics archives (GEO/SRA/BioProject), and the literature (PubMed/OpenAIRE) through one normalized model — deduplicating by DOI, expanding organism queries with NCBI Taxonomy synonyms, and bridging papers to the datasets they produced. Resolves citations and open-access full text, and downloads files.
    6
    4
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server that extracts clean text, tables, and structured data from documents, images, code, and audio files, supporting 97 formats with OCR, transcription, and code intelligence.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables local, read-only extraction of text and structure from PDF, DOCX, PPTX, SVG, and PNG files, including OCR for images, directory tree and metadata reporting, with strict path isolation and audit logging.
    -
  • F
    license
    B
    quality
    A
    maintenance
    Enables local, offline document extraction and manipulation—PDF first but also HTML, DOCX, XLSX, PPTX, EML, EPUB, Markdown, and plain text—through tools for probing, locating, extracting, converting, assembling, OCR, protecting, and redacting documents, with nothing leaving the machine.
    7
    -