model-deploy
Provides tools for inspecting ONNX models, benchmarking ONNX Runtime inference locally or on remote SSH targets, probing deployment environments, and validating deployment constraints such as FPS, P95 latency, and model size.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@model-deploybenchmark ~/models/yolo.onnx on orin-dev and gate at 30 FPS and 35ms P95"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
dsh-model-deploy
Local or remote AI model deployment benchmarking for real hardware targets.
dsh-model-deploy helps answer a practical deployment question: will this model meet my requirements on the machine that will actually run it?
V0.1 focuses on ONNX + ONNX Runtime, with local execution on Windows/Linux and agentless remote execution over SSH to POSIX targets such as Linux workstations and Jetson devices.
V0.1 capabilities
ONNX static inspection: inputs/outputs, shapes, opset, parameter count, operators, model size and dynamic-input detection
Deployment environment probe: OS, CPU, RAM, NVIDIA GPU, ONNX Runtime providers, SSH/SCP
Local ONNX Runtime benchmark: warmup, average latency, P50/P95/P99, FPS
Agentless SSH benchmark: copy a lightweight runner + model to a temporary remote directory, execute, collect JSON, clean up
Deployment Gate: validate measured FPS, P95 latency and model-size constraints
DSH/MCP tools for model inspection, environment probing, local benchmark and remote SSH benchmark
Related MCP server: ComputeSage server
Install
1. Python runtime
The plugin runs its benchmarks in Python. Any interpreter with these packages works:
pip install numpy onnx onnxruntime psutilFor CPU inference, standard onnxruntime is enough. GPU execution requires an ONNX Runtime build/provider compatible with the machine (e.g. onnxruntime-gpu with matching CUDA/cuDNN libraries); if the provider fails to load, results say so rather than reporting CPU numbers as GPU numbers.
The plugin picks the interpreter automatically: DSH_MODEL_DEPLOY_PYTHON if set, then python3/python, then conda/miniforge environments under your home directory — the first one that can import numpy, onnx and onnxruntime.
2. DSH plugin
In DeepSeek Harness, open Plugins → Add plugin and enter:
https://github.com/ethanrise/dsh-model-deploy(or a local checkout path). This registers the model-deploy MCP server and the model-deploy skill; no npm link or PATH setup is needed. Restart DSH after installing or updating.
Standalone CLI (optional)
git clone https://github.com/ethanrise/dsh-model-deploy.git
cd dsh-model-deploy
pip install -e ".[runtime]"
npm install && npm run build # only needed after changing src/
node lib/doctor.js # shows which interpreter was chosen and checks dependenciesCLI
Inspect a model:
dsh-model-deploy inspect model.onnxProbe the current machine:
dsh-model-deploy envBenchmark locally:
dsh-model-deploy bench model.onnx --runs 100Benchmark with deployment constraints:
dsh-model-deploy bench model.onnx \
--min-fps 30 \
--max-p95-ms 35 \
--max-model-mb 200Benchmark a real remote target using an existing SSH alias:
dsh-model-deploy remote-bench orin-dev model.onnx --runs 100The remote target currently needs python3, numpy, and onnxruntime already installed. V0.1 deliberately detects/uses the environment rather than modifying it.
DSH tools
model_inspectdeployment_environmentbenchmark_localssh_preflightbenchmark_remote_ssh
Model paths must be absolute (or start with ~/).
benchmark_local and benchmark_remote_ssh accept inputShapes (e.g. {"images": [1, 3, 640, 640]}) and defaultDynamicDim for dynamic-input models, plus the gate arguments minFps, maxP95Ms, maxModelMb and requireProvider.
Benchmark results report the execution provider the ONNX Runtime session actually used. When ORT silently falls back (e.g. CUDA listed as available but its libraries fail to load), provider_fallback is true and fallback_reason carries ORT's own error. runtime records the Python interpreter, ONNX Runtime and NumPy versions that produced the numbers.
The Python interpreter is chosen from DSH_MODEL_DEPLOY_PYTHON, then python3/python, then conda/miniforge environments under the home directory — the first that can import numpy, onnx and onnxruntime. DSH's MCP client scrubs inherited DSH_* variables, so to pin an interpreter under DSH set it in the patch entry's env.
SSH credentials are not passed to the model. Use ~/.ssh/config, ssh-agent, or the operating system's normal SSH credential handling.
What V0.1 does not do
TensorRT engine building, FP16/INT8 conversion/calibration, Docker/WSL executors, OpenVINO, TFLite, RKNN and NCNN are intentionally deferred. The next major backend target is TensorRT.
Reliability rule
The project distinguishes:
Static analysis — facts read from the model
Compatibility observations — what the detected runtime appears capable of
Measured benchmark — performance actually measured on the named machine
It does not invent latency numbers for hardware that was not tested.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Browser-local and CLI static evidence for deployed AI model artifacts.
Read-only, deterministic AI triage and readiness tools implementing Sophon's published rubrics.
Deterministic validation for AI-generated artifacts: JSON Schema, OpenAPI response, SQL syntax.
Inspect Fireworks AI models, deployments, datasets and fine-tuning jobs.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceAn MCP server that autonomously optimizes ONNX ML models for Arm64 deployment, providing tools to analyze models, apply real INT8 quantization, benchmark performance, and generate Arm64-optimized Docker deployment packages.Apache 2.0
- AlicenseAqualityAmaintenanceEvidence-backed private AI deployment intelligence for GPU and LLM inference. Search benchmark evidence, check deployment fit, predict performance, get hardware recommendations, and generate launch configurations.5GPL 3.0
- AlicenseNot gradedqualityBmaintenanceEnables evaluating AI applications, inspecting reliability evidence, and gating releases from development and CI workflows.30Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables local static analysis of AI deployment artifacts, producing SHA-256 hashes, serialized structure, defects, cautions, evidence gaps, and limitations without executing model code.Apache 2.0