Skip to main content
Glama

clef-use

English | 한국어 | 简体中文 | 日本語

One local computer-use runtime. One MCP interface. A fast decision loop.

clef-use turns a high-level GUI goal into a bounded sequence of visual actions. OmniParser detects screen objects; CLEF selects semantic action candidates; deterministic input adapters execute them. Your planner gets control back at a terminal event instead of making an inference for every click.

flowchart TD
    A[Codex / Hermes / OMP / MCP agent] --> M[Canonical MCP]
    C[Thin CLI] --> R[Shared resident runtime]
    M --> R
    R --> S[Screenshot]
    S --> P[OmniParser V2]
    P --> O[Object map + bounded candidates]
    O --> D[CLEF / CLEF-Flash]
    D --> E[Deterministic OS input]
    E --> V[Verification + progress checks]
    V --> S
    V --> T[Completed / escalation / abort]

Install

The canonical distribution URL is:

curl -fsSL https://ftp.kotori9.dev/clef-use/install.sh | sh

Windows PowerShell:

irm https://ftp.kotori9.dev/clef-use/install.ps1 | iex

After installing the runtime, the installer asks Prepare models now? [Y/n]. Enter or Y runs model preparation; N leaves the runtime installed and skips it. Preparation installs inference dependencies and downloads missing models. Set model_dir to your desired cache path before accepting Y. You can also prepare later:

clef-use models prepare
clef-use doctor

models prepare installs the inference environments, downloads missing model weights, initializes the workers and saves their configuration. Model preparation requires Python 3.11/3.12, Git and sufficient disk space. Set model_dir to your chosen cache or external SSD path before preparation; see installation. Runtime installation supports Python 3.11-3.13 with venv and pip; no root is needed. Updates retain the model cache.

For a source checkout:

uv sync --locked --no-editable
uv run clef-use doctor

Related MCP server: desktop-touch-mcp

Quick start

clef-use models prepare
clef-use doctor
clef-use install-mcp codex
clef-use install-mcp hermes
clef-use install-mcp omp
clef-use run 'In the open Calculator, compute 123 * 456' --success 'Calculator shows 56088'
clef-use status
clef-use abort
clef-use update

Prepare installs pinned ML environments and downloads pinned upstream weights. Use model_dir in configuration for an external SSD or mirror cache. CLEF-Flash alone needs approximately 19.1 GB of weight storage. The larger CLEF needs approximately 55 GB. Cold model loading is separate from inner-loop timing.

MCP

{"mcpServers":{"clef-use":{"command":"clef-use","args":["mcp"]}}}

Five high-level tools: computer_run, computer_continue, computer_observe, computer_status, and computer_abort. No harness-specific executor is required. The CLI and all stdio MCP processes connect to the same authenticated loopback service, sharing loaded models and exclusive desktop input ownership. See MCP and harness setup.

Platforms and verification

V0 targets primary-monitor macOS, Linux and Windows desktops. Windows has a first-class PowerShell installer; model/GUI compatibility outside the measured macOS setup remains experimental. Headless environments support protocol, state-machine, packaging and installer tests. Desktop permissions and an actual GUI are required for real tasks. CUDA, MPS and CPU availability is detected; availability does not establish model compatibility. Measured results and limitations are recorded in evidence.

Benchmarks distinguish fixture overhead from real model/desktop latency. No speedup against conventional VLM computer use is claimed without a controlled comparison. See benchmarking.

Safety

Hard budgets, confidence checks, stale-screen checks, repeated-state detection, explicit cancellation and owned-input release bound execution. Text entry uses exact supplied or goal-extracted literals. Password-like targets are refused. Constraints are model-assisted; this runtime is not an OS sandbox. Screen content can be malicious, and a decision model can make mistakes. Use an isolated desktop with non-sensitive tasks. Screenshots stay local unless a planner explicitly requests an MCP image. No shell execution tool is exposed.

Documentation

License

Runtime code: Apache-2.0. Model code, weights and dependencies retain their upstream licenses. The pinned OmniParser YOLOv9 detector uses MIT-licensed code; earlier Ultralytics detectors carry different licensing. See third-party notices.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI clients to automate Windows desktop applications through window manipulation, image recognition, OCR, keyboard/mouse simulation, and memory operations via the MCP protocol.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI clients on Windows to control the local mouse, keyboard, and screen understanding via MCP stdio, allowing automated workflows like viewing the screen, locating elements, clicking, typing, and verifying results.
    1
    Apache 2.0