Skip to main content
Glama
jiaxi1102
by jiaxi1102

o2-mcp

Generic, project-agnostic access to the HMS O2 cluster, exposed both as a Python library (o2mcp) and as an MCP server (o2-mcp) so an agent can submit Slurm work, run remote commands, monitor jobs, move files, and keep disk tidy — without triggering a Duo push on every action.

Extracted from clock-oscillation-analysis so the cluster tooling is shared infrastructure (used by multiple analysis projects) rather than living inside one of them. Project-specific layers (e.g. run-organization for a particular pipeline) build on this package rather than living in it.

Duo model (read this first)

Our August 2026 incident testing showed that HMS O2 can issue Duo challenges not only for a new SSH transport, but also when a new session channel is opened inside an existing OpenSSH ControlMaster. A normal ssh o2 command invocation therefore is not Duo-safe merely because it reaches the expected mux socket.

Version 0.4 replaces that command pattern with one workstation-wide broker per configured host role. Login-node work normally needs only the login broker; a separate transfer broker exists for commands that must execute on the transfer host:

  1. o2_start_broker consumes a short-lived, client-bound, one-attempt role-matched grant and starts exactly one SSH process for that broker. When no grant is supplied, the MCP may mint that same one-attempt grant automatically only after proving the selected O2 host routes through the HMS VPN.

  2. That SSH process opens exactly one remote session running a small embedded Python helper.

  3. Every later o2_exec, Slurm, workspace, keepalive, and run-organization command is length-prefixed JSON sent through the same session's stdin/stdout.

  4. Independently launched MCP tasks share the applicable broker through a mode-0600 Unix socket. Commands are serialized within each role in the MVP.

  5. The daemon never reconnects. If the channel dies, its socket disappears and later commands fail locally; another login requires another explicit grant.

The workstation-wide ~/.agent_locks/O2_POLICY.json still defaults fail-closed. Both the MCP client and the broker daemon check it before every logical command, so a concurrent disabled transition wins even against a hand-crafted local socket request. Disabling does not terminate a command already in progress.

The broker uses an O2B1 magic marker, four-byte network-order lengths, and UTF-8 JSON rather than newline framing. Remote stdout/stderr is drained while retaining at most 1 MiB per stream, so noisy commands cannot grow broker memory without bound and newlines or JSON-looking output cannot corrupt the next command. Frames are limited to 16 MiB, while command text is capped at 64 KiB so bash -c stays below the remote execve argument limit. One command timeout returns code 124 without reconnecting or replacing the persistent channel. Finite deadlines are capped at seven days to avoid platform socket overflow; an explicit None retains indefinite execution. Logical commands use a non-login Bash that inherits the one session environment, avoiding repeated profile banners and startup latency. Full escaped requests are size-checked before dispatch, stdin is capped at 1 MiB, and transport writes use inactivity plus frame-size-scaled deadlines so a slow progressing link survives while a stall cannot retain the global policy mutex indefinitely. For finite commands, the daemon also bounds the result-frame wait with an inactivity timeout and a frame-size-scaled absolute budget; slow progress survives, while a silent helper terminates and unpublishes the one transport instead of leaving a falsely reusable broker wedged.

ControlMaster hardening remains in the library for the transfer compatibility layer and offline regression tests. It is not the login command boundary: both the MCP wrapper and public O2Connection.start_master() API reject login master starts and raw SSH, directing callers to the broker instead. See the broker design.

OpenSSH evaluates Match exec shell predicates even for its nominally local ssh -G config dump. The MCP therefore reads the configured SSH file and recursively flattens Include directives itself, rejects any Match block before launching OpenSSH, and runs ssh -G only against that private inspected snapshot. If your normal SSH config contains Match, put the O2 Host blocks in a separate Match-free file and set O2_SSH_CONFIG_FILE to it.

The optional o2-transfer alias remains a different host with its own grant, command broker, and rsync-compatibility ControlMaster. Run transitions and explicit transfer-host probes use the framed transfer broker. Existing detached rsync transfers remain on their pinned ControlMaster path and are preserved; that compatibility path must not be described as Duo-free merely because its ControlMaster exists.

Never open the broker/master in a loop or run authentication tools on a timer. The policy file has two durable modes: disabled blocks new remote operations, while reuse_only permits only existing exact sockets. There is deliberately no durable normal mode. The one-shot grant is atomically removed and converted to an active attempt receipt before SSH, so a failure, timeout, or crash cannot leave reusable authorization for a queued task. Policy generation/revision pairs and a stable internal mutex serialize every transition across MCP processes. Callers pass both values from one o2_local_status snapshot; repair creates a new generation, so a repeated numeric revision cannot revive stale approval.

If repeated Duo prompts begin while the native MCP tools are unavailable, use the package's local-only emergency command. It performs the same atomic policy transition as o2_policy_disable and never invokes SSH or any remote probe:

o2-mcp-policy-disable --reason "repeated Duo prompts"

This command can only reduce authority. Re-enabling reuse and issuing a login grant remain MCP-only operations with explicit global or target-scoped approval.

Be on the HMS VPN. O2 normally skips Duo for connections from HMS-trusted source IPs — i.e. when your SSH egresses through the HMS VPN (GlobalProtect), not your normal internet interface. If the VPN is down (or split-tunnel isn't routing O2's subnet), the one o2_start_broker login or transfer-host startup may Duo-push. To make accidental off-VPN launch impossible, o2_start_broker and the transfer-only o2_start_master refuse to open a new login unless the route to the selected O2 host is specifically bound to HMS GlobalProtect. The guard checks route get locally, reads the selected interface's live address, and requires that address to match a PanGPS-owned preferred tunnel IP for the configured HMS portal. This prevents an unrelated VPN's utun* interface from being mistaken for GlobalProtect; all checks are local (no connection, no Duo). An explicitly approved grant may scope allow_offvpn: true to that one attempt. There is no durable bypass. The defaults can be pointed at another managed HMS installation with O2_GLOBALPROTECT_SETTINGS_FILE, O2_GLOBALPROTECT_PORTAL, and O2_VPN_IFACE_PREFIX.

The start tools default auto_authorize_on_vpn to true. If the requested broker/master is absent and the route is proven on-VPN, the tool records the standing authorization in the policy audit log as standing_on_vpn (distinct from explicit_user_approval), mints one allow_offvpn=false grant, consumes it for one start, and never retries. If routing is off-VPN or indeterminate, it returns error: off_vpn before opening SSH; the agent must ask the user before separately issuing an allow_offvpn=true grant. This automatic connection permission does not authorize transfer, submission, cancellation, deletion, promotion, archive, or any other remote mutation.

Related MCP server: tacc-mcp-bio

Install

# The core (config/connection/sync/slurm/async_transfer/keepalive/workspace) is pure-stdlib
# and runs on Python 3.9. The MCP server needs the mcp SDK (Python >= 3.10):
pip install -e ".[o2]"     # on a 3.10+ env

MCP server config

{
  "mcpServers": {
    "o2": {
      "type": "stdio",
      "command": "/path/to/venv/bin/o2-mcp",
      "env": {
        "O2_SSH_HOST_ALIAS": "o2",
        "O2_SSH_TRANSFER_ALIAS": "o2-transfer",
        "O2_SSH_CONFIG_FILE": "/Users/you/.ssh/config",
        "O2_POLICY_FILE": "/Users/you/.agent_locks/O2_POLICY.json",
        "O2_BROKER_DIR": "/Users/you/.agent_locks/o2-broker",
        "O2_TRANSFER_BROKER_DIR": "/Users/you/.agent_locks/o2-transfer-broker"
      }
    }
  }
}

Requires a Host o2 block (and optionally Host o2-transfer) in the inspected SSH config. The broker snapshots that Match-free config before consuming the grant and explicitly disables mux reuse, proxies, and local command hooks for its one transport. O2_BROKER_DIR must be absolute, private, and short enough for a macOS Unix socket; the default satisfies those constraints. O2_TRANSFER_BROKER_DIR has the same requirements and must be distinct from O2_BROKER_DIR.

Tools

Tool

Purpose

Hint

o2_local_status

Local policy, sockets, processes, receipts, and transfer logs; never SSH

read-only/local

o2_status

Deprecated local-only compatibility alias

read-only/local

o2_policy_disable

Block every new remote O2 operation without killing existing processes

write/local

o2_policy_enable_reuse

Explicitly enable existing broker/transport reuse at an observed generation/revision

write/local

o2_authorize_login

Issue one short-lived login or transfer-host grant

write/local

o2_start_broker

Reuse a broker or open one role-specific command channel; without grant_id, auto-authorize one attempt only over a proven HMS VPN route (transfer=true for the transfer host)

write

o2_stop_broker

Locally close one role-specific broker and its SSH process; no policy change

destructive/local

o2_start_master

Transfer compatibility only; without grant_id, auto-authorize one attempt only over a proven HMS VPN route; login-master starts are rejected

write

o2_stop_master

Locally close the transfer master or a pre-upgrade login master through its exact socket; shutdown-only and does not restore login startup

destructive/local

o2_probe

One explicit fixed command through the existing broker; never retried

read-only

o2_exec

Run an arbitrary command on a login node

write

o2_submit_job

sbatch a script (existing path or staged script_text); returns the job id

write

o2_squeue

squeue -u <user> as structured rows

read-only

o2_job_status

sacct -j <id> accounting (state, elapsed, exit code, MaxRSS)

read-only

o2_tail_log

Tail a remote log file

read-only

o2_cancel_job

scancel <id>

destructive

o2_push / o2_pull

rsync up/down through the existing transfer master by default; use_transfer_node=false is legacy reuse only

write

o2_push_async / o2_pull_async

Non-blocking rsync: launch detached, return a transfer_id immediately

write

o2_transfer_status

Progress/state of async transfers (running/done/failed/crashed); omit id to list all

read-only

o2_transfer_cancel

SIGTERM a running async transfer's process group

destructive

o2_disk_report

Per-tier usage + hygiene flags (regenerable/redundant/misplaced)

read-only

o2_workspace_gc

Prune regenerable + redundant disk (detached, dry-run default)

destructive

o2_place

Resolve the canonical output path for a kind (+project) per tier

read-only

Authenticated execution-plan contract

o2mcp.runorg includes a project-neutral, immutable ExecutionPlan contract for pipelines that need stronger provenance than a standalone sbatch script. A plan binds:

  • project, campaign, pipeline, registered run, source-commit, and source-bundle identities;

  • dataset IDs, reviewed manifest digests, and optional independently generated storage-binding digests;

  • absolute run/work/results/receipt/log/promotion paths;

  • an acyclic stage graph, signed afterok/afterany dependency semantics, optional stable array tasks, and exact per-stage or per-task argv vectors;

  • absolute executables and working directories plus a fingerprint of the exact executable (normally an immutable environment-validating wrapper), and a clean environment containing only explicit non-secret values;

  • bounded structured Slurm resources (including GPU model, constraints, exclusions, and licenses) and retry policies; and

  • every stage- and task-level receipt expected below the run's canonical receipt tree.

Plans are frozen dataclass object graphs. ExecutionPlan.to_json() writes a self-checksummed envelope whose plan_sha256 is derived from compact, canonical JSON; ExecutionPlan.from_json() rejects digest drift, duplicate JSON members, unsupported fields, malformed paths, unbounded retries, duplicate receipts, and cyclic or incomplete dependencies. Supplying the independently reviewed digest as expected_plan_sha256 authenticates the plan at a submission boundary; the unkeyed digest embedded in the file is only an integrity check. Set-like collections are canonicalized before hashing, so renderer insertion order cannot change the idempotency key.

This contract is intentionally execution-only: GEM, clock, and future pipeline repositories still render their scientific commands and define their own QC and release criteria.

Idempotent execution and reconciliation

PreparedRunIdentity allocates the run ID/root and dataset scope before those values are sealed into an ExecutionPlan. ExecutionEngine then submits stages through a fakeable ExecutionBackend with these guarantees:

  • production execution authenticates the plan against the prepared active run.json—including project and exact canonical path—before writing the plan binding or contacting Slurm;

  • before calling sbatch, it writes an immutable submission intent containing the plan SHA, stage, attempt, selected stable tasks, script SHA, resources, and dependency arguments, then atomically publishes a separate invocation marker;

  • Slurm receives a reversible o2plan:v1:<plan-sha>:<stage>:a<attempt> comment. Repeated calls and lost responses query this identity. An intent whose creator died before invocation is safely recoverable by the invocation-marker winner. Once invocation is owned, calls remain query-only until the original job becomes visible; there is deliberately no timeout-based takeover that could duplicate a late job;

  • successful array tasks receive immutable per-attempt receipts and are excluded from bounded missing-only retries. An explicitly terminal CANCELLED, TIMEOUT, or FAILED task is not reinterpreted as missing simply because its output is absent; it retries only when the signed state/exit policy permits it;

  • receipt probes distinguish authoritative absence from transport/parser failure. Untrustworthy observations remain transient and cannot authorize a retry or be frozen into task evidence;

  • current execution state is synchronized to run.json and the JSONL registry. A registry outage leaves a coalesced pending update that can be replayed without touching Slurm; and

  • scheduler options are rendered from typed plan fields. Arbitrary arguments cannot override the plan SHA comment, array selection, dependency, resources, working directory, or logs.

Dependency-bound reconciliation is explicit in the signed DAG. A pipeline adds a non-array reconciler StageSpec whose dependencies are the compute stages and whose dependency_mode is afterany; its signed command performs the pipeline's files-as-truth validation. ExecutionEngine.submit_afterany_reconciler() then submits that command with --dependency=afterany:<latest-job-ids>, so it runs after every prerequisite reaches a terminal state even when an array is partial. reconcile_stage() is also available for workstation recovery/observability and for launching policy-bounded missing-only attempts, but caller polling is not the dependency mechanism.

Reconciliation JSON is decoded as a strict plan/stage/attempt contract and must cover the exact task set with matching immutable task-attempt evidence before it can release an afterok stage. Each accepted missing-only retry also creates a durable follow-up outbox authorization. Replaying the accepted retry or any later reconciliation converges the same afterany generation after a coordinator crash, without resubmitting completed compute work.

Non-blocking transfers

o2_push_async / o2_pull_async launch a detached rsync and return a transfer_id right away, so the agent can keep working and poll o2_transfer_status instead of blocking a tool call for a multi-GB transfer. The transfer keeps running between tool calls and survives an MCP-server restart (a wrapper records rsync's exit code to disk); re-running the same command resumes it (rsync --partial). Remote paths are escaped so spaces transfer intact while ~/$VAR/${VAR} still expand. State lives under ~/.cache/clock_o2_mcp/transfers (O2_ASYNC_STATE_DIR to override).

Safety contract

  • O2_POLICY.json is the sole policy state. Missing, malformed, symlinked, wrong-owner, or permissively readable state is effectively disabled; no project/ancestor lock files or bypass environment variables are consulted. O2_POLICY_FILE must be absolute. The first explicit policy mutation safely tightens an owned physical legacy ~/.agent_locks directory from 0755 to 0700; aliased or foreign-owned directories remain invalid.

  • Only o2_start_broker (role-specific command transport) or the transfer rsync compatibility start can authenticate, and only after atomically consuming a matching client/host/off-VPN-scoped grant. The parent starts only a local daemon under the policy mutex; that daemon revalidates the exact active attempt and holds the mutex around its sole SSH spawn. Launch data travels only through a bounded inherited anonymous pipe, so another same-UID task cannot replace it through the filesystem or replay a durable recipe.

  • Login and transfer-host commands never invoke SSH directly and never open a new ControlMaster session channel. Raw SSH through run_raw is rejected in production. Rsync remains the explicitly isolated compatibility boundary.

  • Each broker socket directory is physical, caller-owned, and mode 0700; its socket and files are mode 0600. One lifetime flock per role prevents overlapping daemons. An unbindably long socket path fails before grant consumption. Clients require a trusted ready receipt for the exact protocol version, role alias, and expanded HostName/User/Port; incomplete local frames expire on a finite absolute deadline. Destination changes block command reuse but not local broker stop, so the stale daemon can be retired without contacting O2.

  • The library retains its historical require_master parameters for source compatibility, but rejects require_master=False; callers cannot opt back into cold SSH/rsync behavior.

  • Blocking and detached rsync default to the transfer alias, whose master can be started with a transfer-scoped one-shot grant. Login-alias rsync remains only for reusing an already-existing legacy login master; the MCP cannot create one.

  • o2_local_status and the deprecated o2_status alias inspect broker receipts, Unix sockets, processes, policy, and transfer logs locally; neither invokes SSH. o2_probe is the only status-like remote operation and runs exactly once through the selected role-specific broker.

  • The policy JSON contains the login-attempt receipt and five-minute cooldown. A zero exit from ssh -MNf is successful only when an immediate exact-socket control check confirms that the background master survived.

  • disabled does not automatically stop a master or detached transfer. Local inspection and separately approved local transfer cancellation remain available.

  • Destructive/transfer-node operations default to dry-run where applicable and verify before freeing scratch.

Development

pip install -e ".[dev,o2]"
ruff check src tests && black --check src tests && pytest -m "not o2" -q

The core stays import-light (stdlib only); mcp/pydantic/anyio are needed only by the server. Tests inject the subprocess seam, so they run fully offline (no cluster, no network).

F
license - not found
Not graded
quality - not tested
B
maintenance

Maintenance

0Releases (12mo)
Commit activity
Issues opened vs closed

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables Claude Code to interact with a TACC or SLURM HPC cluster for bioinformatics pipelines, allowing job management, log reading, file browsing, remote script execution, and job submission through natural language.
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to manage SLURM HPC clusters via SSH. Supports job submission, resource monitoring, queue management, and file operations.
    14
    4
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables executing commands on remote SSH hosts, with full support for bastion/jump hosts and ~/.ssh/config, plus Slurm job management and rsync.
    3
    MIT

View all related MCP servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/jiaxi1102/o2-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server