Skip to main content
Glama

omlx-mcp-server

Local stdio MCP server for oMLX.

Tools

  • omlx_status: return the local oMLX model list and current power status in one call.

  • omlx_run: unified execution tool for either chat or agent mode.

Related MCP server: local-mcp

Battery Guard

omlx_run checks power status before using a model.

  • On AC power: execute normally.

  • On battery power: return status="needs_confirmation" unless allow_on_battery=true.

This is meant to force an explicit user decision before running heavy local inference on battery.

Compact Interface

The server is intentionally compressed to two tools to keep MCP schema overhead down.

  • Use omlx_status() to fetch models + power_status.

  • Use omlx_run(mode="chat" | "agent", prompt=...) for execution.

omlx_run picks a default model automatically:

  • chat mode defaults to MLX-Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-4bit

  • agent mode defaults to MLX-Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-8bit

Defaults

  • Default chat model: MLX-Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-4bit

  • Default agent model: MLX-Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-8bit

  • Default base URL: http://127.0.0.1:8000/v1

Local Run

UV_CACHE_DIR=.uv-cache UV_PROJECT_ENVIRONMENT=.venv uv sync --dev
UV_CACHE_DIR=.uv-cache UV_PROJECT_ENVIRONMENT=.venv uv run omlx-mcp-server

Codex Config Snippet

Add this to your Codex config if you want future sessions to discover it automatically. Replace /path/to/omlx-mcp-server with your clone path and set your local oMLX key:

[mcp_servers.omlx]
command = "uv"
args = [
  "run",
  "--directory", "/path/to/omlx-mcp-server",
  "omlx-mcp-server",
]

[mcp_servers.omlx.env]
UV_CACHE_DIR = "/path/to/omlx-mcp-server/.uv-cache"
UV_PROJECT_ENVIRONMENT = "/path/to/omlx-mcp-server/.venv"
OMLX_BASE_URL = "http://127.0.0.1:8000/v1"
OMLX_API_KEY = "your-local-omlx-key"
OMLX_DEFAULT_MODEL = "MLX-Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-4bit"
OMLX_AGENT_MODEL = "MLX-Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-8bit"

Available Tools

2 tools
omlx_runB

Run either a chat or agent-style local oMLX request. Blocks on battery unless allow_on_battery is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYes
modelNo
promptYes
contextNo
max_tokensNo
temperatureNo
system_promptNo
allow_on_batteryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does disclose one important behavior: 'Blocks on battery unless allow_on_battery is true,' which adds context about conditional execution. However, it does not describe other relevant behaviors such as output format, error handling, or side effects, leaving significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that communicates the core purpose and a key behavioral constraint without extraneous words. It is front-loaded and easy to parse, with no wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 8 parameters, 2 required, and no annotation coverage, the one-sentence description is inadequate. It does not explain the meaning of key parameters (mode, prompt, context, etc.), nor does it set expectations for return values despite an output schema existing. The description is too sparse for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only explains the semantics of one parameter (allow_on_battery) by tying it to the blocking behavior. The other seven parameters (mode, model, prompt, context, max_tokens, temperature, system_prompt) are left entirely to the schema, which provides no descriptions. This is insufficient for an 8-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Run either a chat or agent-style local oMLX request,' identifying both the verb and the resource. It distinguishes between two modes (chat/agent), which adds specificity. However, it does not explicitly contrast with the sibling tool omlx_status, so it misses a clear differentiation opportunity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor does it mention the sibling tool omlx_status. It only implies usage by describing what it does. There is no indication of appropriate contexts, prerequisites, or exclusions, leaving the agent to infer when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omlx_statusA

Return local oMLX model list and current Mac power status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. The verb 'Return' implies a read-only operation, but the description does not explicitly state that no modifications are made or add other behavioral context such as scope or side effects. For a simple status tool, this is adequate but not fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is direct and free of unnecessary words. It earns its place by stating exactly what the tool does without redundancy or clutter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 parameters, output schema available), the description is complete. It clearly states what is returned, and the presence of an output schema covers return value details. No further context is needed for an agent to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description need not explain parameter meaning. According to the rubric, a baseline of 4 applies when there are no parameters, and the description does not need to compensate for schema gaps since none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Return' and clearly identifies the two resources: local oMLX model list and current Mac power status. This makes the tool's purpose unambiguous and distinguishes it from the sibling tool omlx_run, which presumably runs models.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when one wants to inspect models or power status, but it does not explicitly state when to use this tool versus omlx_run, nor does it mention any exclusions or prerequisites. It provides context but no direct guidance on alternative selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv0.1.0
    • First observedomlx_run
    • First observedomlx_status

TDQS

A3.7/5.0
Disambiguation5/5

The two tools have completely distinct purposes: one for querying status/model list, the other for running requests. There is no overlap or ambiguity.

Naming Consistency5/5

Both tools follow a consistent 'omlx_' prefix plus a verb (status, run). This is a clear and predictable pattern.

Tool Count3/5

With only 2 tools, the set feels thin for a model-running server. While each tool is justified, the count is on the borderline of being too minimal.

Completeness4/5

The core functionality is covered: status provides information needed before running, and run executes the request. Minor gaps exist (e.g., no explicit model management or request cancellation), but for the stated purpose the surface is largely complete.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    A Python MCP server that exposes local Ollama models as tools for AI assistants, enabling chat, generation, embeddings, and model management without cloud APIs.
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables running MCP tools against local MLX models on your Mac, with hardware-aware configuration, CLI streaming, and a dashboard for routing and monitoring.
    3,443
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server that enables local Apple on-device Foundation Model access via any MCP client, supporting text generation, structured output, and multi-turn chat on macOS.
    2
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    MCP server that connects LLM agents to a local LM Studio instance, enabling model management, OpenAI-compatible chat completions, text completions, and embeddings through a set of tools.
    9
    1
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/glasses666/omlx-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server