clef-use
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@clef-useIn the open Calculator, compute 123 * 456 and verify it shows 56088."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
clef-use
One local computer-use runtime. One MCP interface. A fast decision loop.
clef-use turns a high-level GUI goal into a bounded sequence of visual actions. OmniParser detects screen objects; CLEF selects semantic action candidates; deterministic input adapters execute them. Your planner gets control back at a terminal event instead of making an inference for every click.
flowchart TD
A[Codex / Hermes / OMP / MCP agent] --> M[Canonical MCP]
C[Thin CLI] --> R[Shared resident runtime]
M --> R
R --> S[Screenshot]
S --> P[OmniParser V2]
P --> O[Object map + bounded candidates]
O --> D[CLEF / CLEF-Flash]
D --> E[Deterministic OS input]
E --> V[Verification + progress checks]
V --> S
V --> T[Completed / escalation / abort]Install
The canonical distribution URL is:
curl -fsSL https://ftp.kotori9.dev/clef-use/install.sh | shWindows PowerShell:
irm https://ftp.kotori9.dev/clef-use/install.ps1 | iexAfter installing the runtime, the installer asks Prepare models now? [Y/n].
Enter or Y runs model preparation; N leaves the runtime installed and skips it.
Preparation installs inference dependencies and downloads missing models. Set
model_dir to your desired cache path before accepting Y. You can also prepare later:
clef-use models prepare
clef-use doctormodels prepare installs the inference environments, downloads missing model
weights, initializes the workers and saves their configuration. Model preparation
requires Python 3.11/3.12, Git and sufficient disk space. Set model_dir to your
chosen cache or external SSD path before preparation; see installation.
Runtime installation supports Python 3.11-3.13 with venv and pip; no root is
needed. Updates retain the model cache.
For a source checkout:
uv sync --locked --no-editable
uv run clef-use doctorRelated MCP server: desktop-touch-mcp
Quick start
clef-use models prepare
clef-use doctor
clef-use install-mcp codex
clef-use install-mcp hermes
clef-use install-mcp omp
clef-use run 'In the open Calculator, compute 123 * 456' --success 'Calculator shows 56088'
clef-use status
clef-use abort
clef-use updatePrepare installs pinned ML environments and downloads pinned upstream weights.
Use model_dir in configuration for an external SSD or mirror cache.
CLEF-Flash alone needs approximately 19.1 GB of weight storage. The larger CLEF
needs approximately 55 GB. Cold model loading is separate from inner-loop timing.
MCP
{"mcpServers":{"clef-use":{"command":"clef-use","args":["mcp"]}}}Five high-level tools: computer_run, computer_continue, computer_observe,
computer_status, and computer_abort. No harness-specific executor is required.
The CLI and all stdio MCP processes connect to the same authenticated loopback
service, sharing loaded models and exclusive desktop input ownership.
See MCP and harness setup.
Platforms and verification
V0 targets primary-monitor macOS, Linux and Windows desktops. Windows has a first-class PowerShell installer; model/GUI compatibility outside the measured macOS setup remains experimental. Headless environments support protocol, state-machine, packaging and installer tests. Desktop permissions and an actual GUI are required for real tasks. CUDA, MPS and CPU availability is detected; availability does not establish model compatibility. Measured results and limitations are recorded in evidence.
Benchmarks distinguish fixture overhead from real model/desktop latency. No speedup against conventional VLM computer use is claimed without a controlled comparison. See benchmarking.
Safety
Hard budgets, confidence checks, stale-screen checks, repeated-state detection, explicit cancellation and owned-input release bound execution. Text entry uses exact supplied or goal-extracted literals. Password-like targets are refused. Constraints are model-assisted; this runtime is not an OS sandbox. Screen content can be malicious, and a decision model can make mistakes. Use an isolated desktop with non-sensitive tasks. Screenshots stay local unless a planner explicitly requests an MCP image. No shell execution tool is exposed.
Documentation
English setup, 한국어, 简体中文, 日本語
License
Runtime code: Apache-2.0. Model code, weights and dependencies retain their upstream licenses. The pinned OmniParser YOLOv9 detector uses MIT-licensed code; earlier Ultralytics detectors carry different licensing. See third-party notices.
This server cannot be deployed
Maintenance
Related MCP Connectors
Hosted MCP memory and agent control plane for durable conversations, jobs, and operations.
Automate 1,000+ services from any MCP-compatible AI agent: build Applets, run actions and queries.
The governed runtime for agent skills. Search the catalog and inspect a skill before running it.
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI clients to automate Windows desktop applications through window manipulation, image recognition, OCR, keyboard/mouse simulation, and memory operations via the MCP protocol.MIT
- AlicenseAqualityCmaintenanceAllows AI clients to see and control Windows 10/11 desktops via MCP, with screenshots, UI Automation, Chrome CDP, keyboard/mouse, and terminal using semantic element targeting.30541 npmMIT
- AlicenseAqualityAmaintenanceEnables MCP agents to automate real GUI applications on headless desktops, providing background mouse/keyboard control, window/process management, screenshots, and safe human handoff without disturbing the user's desktop.584MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI clients on Windows to control the local mouse, keyboard, and screen understanding via MCP stdio, allowing automated workflows like viewing the screen, locating elements, clicking, typing, and verifying results.1Apache 2.0