Skip to main content
Glama

Oh My Laya

English | 中文

One command to bring local Laya decision-making to your coding agents.

Oh My Laya builds on laya-mlx. It downloads and verifies Hugging Face weights, then registers the laya_tell_me tool for classification, scoring, risk routing, and yes/no decisions. Inference runs entirely on your Mac after the model is downloaded.

It supports Codex, Claude Code, DeepSeek Harness (DSH), and pi-agent. The installer detects available clients and lets you select one, several, or all of them.

Install

Requirements: Apple Silicon, macOS 14+, and Python 3.11+. The default multilingual FP16 checkpoint is approximately 678 MB.

Install and register with all detected clients:

sh -c "$(curl -fsSL https://raw.githubusercontent.com/leo1394/oh-my-laya/master/tools/oh-my-laya.sh)"

Without curl, use wget:

sh -c "$(wget -qO- https://raw.githubusercontent.com/leo1394/oh-my-laya/master/tools/oh-my-laya.sh)"

Install your preferred agent client first; Oh My Laya connects to clients already on your machine.

Already cloned the repository? Run the interactive installer from its root:

./install.sh

For non-interactive installation:

./install.sh --targets codex
./install.sh --targets codex,claude
./install.sh --targets all

Both all and both register every detected client. Restart the affected agent sessions after installation.

For Codex, installation adds Oh My Laya as a local plugin with its logo and MCP tools. Open Plugins → Personal → Oh My Laya, then start a new task to use it. An up-to-date Codex CLI with plugin add support is required. Re-run the same installer to update. Laya Model Advisor and Alpha Squad remain separately installed Skills; Alpha Squad is fetched from GitHub.

Start the local Snake demo after installation:

laya --snake

Demo dependencies and the installed model are configured automatically. Use laya --snake --help for demo options. If the installer reports that the command directory is missing from PATH, follow its displayed PATH instruction once.

For each selected client, the installer asks Use Alpha Squad + Laya with /goal? [y/N]. Choose y to add the Goal workflow to that client's global instructions, using actual installed skill paths. Existing Goal rules are backed up and replaced; other sections are preserved. Choose n to leave global instructions unchanged.

For unattended installation, choose explicitly:

./install.sh --targets all --goal-workflow yes
# Or leave global instructions unchanged:
./install.sh --targets all --goal-workflow no

Client

Default global instructions

Codex

~/.codex/AGENTS.md

Claude Code

~/.claude/CLAUDE.md

DSH

~/.dsh/AGENTS.md

pi-agent

~/.pi/agent/AGENTS.md

Custom client home directories are respected. After opting in, start a new session and use /goal for a task. Codex, Claude Code and DSH keep their native Goal implementation (the installed version/profile must support it). Pi gets a /goal prompt template, not a persistent Goal loop; conflicting custom prompts are preserved and reported. Model selection and delegation require the host's verified capabilities; unavailable features need an explicit manual fallback. Without an interactive terminal, global rules are unchanged unless explicitly opted in.

Related MCP server: local-mcp

Use It

Start with a goal

Enabled /goal integration during installation? Open a new Codex session, enter Goal mode with /goal, and give Codex a task:

Add search to this project, cover it with tests, and review the changes.
  1. Confirm your strategy. On first setup, one window collects your Laya policy, execution model and reasoning ceiling, reviewer configuration, and final Confirm and continue. Choose automatic recommendations for hands-off subagent routing; saved settings are revalidated on later tasks.

  2. Let Codex split the work. Alpha Squad coordinates exploration, implementation, testing, research, and review as needed. Small tasks can stay with the main agent.

  3. Let Laya guide subagent assignments. Laya assesses the work; Alpha Squad applies accepted model/effort recommendations through the host. Automatic execution routing uses only your selected model, at or below your reasoning ceiling. The main session's model stays unchanged.

  4. Complete and review together. Subagents return focused results to the main agent, which integrates and verifies the work. Difficult, high-risk, or uncertain reviews use the main session's model/effort; ordinary reviews use your configured reviewer.

flowchart TD
    goal["/goal + your task"] --> setup["Confirm policy, models and ceilings"]
    setup --> lead["Orchestrator: current model unchanged"]
    lead --> split{"Delegate useful subtasks?"}
    split -->|"No"| solo["Main agent handles task"]
    split -->|"Yes"| laya["Local Laya: assess and recommend"]
    laya --> route["Alpha Squad: apply policy and verified assignments"]
    subgraph squad ["Execution roles — only as needed, within your model / effort ceiling"]
        explorer["Explorer: map code"]
        researcher["Researcher: verify facts"]
        worker["Worker: implement"]
        tester["Tester: validate"]
    end
    route --> explorer & researcher & worker & tester
    explorer & researcher & worker & tester --> merge["Orchestrator: integrate results"]
    merge --> review["Reviewer: configured model; main model for difficult reviews"]
    review --> verify["Orchestrator: verify and deliver"]
    solo --> verify

Roles show responsibilities, not a fixed parallel schedule. Alpha Squad orders dependent work and requests confirmation when your policy requires it.

Designed to reduce token waste: local Laya handles lightweight routing decisions, while focused subtasks and bounded reasoning help avoid unnecessary large-model work. Actual token savings depend on the task; subagent coordination adds overhead, so delegation is used only when useful. Model choice alone does not guarantee fewer tokens.

The confirmation window and model assignments require host support. Execution approvals still apply; this workflow does not grant permission for sensitive actions.

Ask Laya directly

For a standalone decision without the squad workflow:

Use Laya to classify the current change as low, medium, or high risk, and decide whether human review is needed.

Oh My Laya provides these tools:

Tool

Purpose

laya_tell_me

choice classification, score ranking, and noul yes/no probability

laya_advisor_preferences

Read or change model recommendation preferences

Laya does not generate code and must not authorize destructive, publishing, or other consequential actions. The multilingual model has a total context budget of 1,024 tokens, so ask the agent to summarize long inputs first.

Codex Model Advisor

Each selected client gets the advisor skill and the latest Alpha Squad from GitHub. Existing unmanaged or locally modified Alpha Squad installations are preserved. Restart the client. In Codex, ask:

Use $laya-model-advisor for each new task in this session. Ask me to choose a recommendation policy.

Choose always ask, ask only for high complexity, high risk or uncertainty, or automatically accept recommendations. You can change this preference in chat at any time; it persists across sessions. When asking, the advisor lists verified available models and lets you choose a model and its supported reasoning effort. Without a user-defined model tier mapping, it recommends effort changes on the current model; you can still choose another listed model.

This is advisory: accepting a recommendation does not switch the active model. Apply it in Codex's model picker. Popups require host support; otherwise the agent asks in chat. Session advice is skill-driven, not a guaranteed per-message hook. If the model list cannot be verified, the advisor asks you to provide it. These preferences never bypass execution approvals.

Setup uses one window: policy → model → reasoning → Confirm and continue. In automatic mode, model and effort are still required: advice stays on the selected model and never exceeds the selected effort ceiling (proposed default: high). Locally disabled efforts remain hidden. Existing automatic preferences without a ceiling require setup again.

Route subagent models

Use $alpha-squad-coding-craft with $laya-model-advisor. Configure Laya-based subagent routing.

One window configures policy, execution model/effort, reviewer model/effort, and Confirm and continue. The main session stays unchanged. Execution agents use Laya's accepted recommendations; auto stays within your chosen model and effort ceiling. Difficult, high-risk or uncertain reviews use the main session's exact model/effort; ordinary reviews use the configured reviewer. Alpha Squad assigns subagent models through the host, not through a main-session model switch.

Without Laya, Alpha Squad still works with manual model selection in one window. Rerun the installer to fetch upstream updates; local customizations are never silently overwritten. A network failure is reported, not treated as an update.

Common Options

# Select another checkpoint
./install.sh --targets all --model english
./install.sh --targets dsh --model typed-decisions

# Preview without downloading or changing configuration
./install.sh --targets codex --dry-run

The default installation directory is ~/.local/share/oh-my-laya/. Each agent starts its own lazy MCP process; concurrent callers each consume a separate unified-memory allocation.

Development

PYTHONPATH=src python3 -m unittest discover -s tests -v

Licensed under the MIT License.

Available Tools

1 tool
laya_tell_meB

Run a local typed decision with Laya-MLX.

Use for bounded classification, ordered scoring, and yes/no probability. Question types are choice, score, and noul. Treat results as advisory; never use them as authorization for destructive or consequential actions.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes
questionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that results are advisory and not authorization for destructive actions, which is valuable behavioral context. However, it doesn't describe any other behavioral traits such as potential side effects, error handling, or whether the operation is synchronous or has latency. The advisory warning is a positive, but it stops short of a comprehensive disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the core action, the second gives usage scope, the third is a safety caveat. No filler or repetition. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given two parameters, nested objects, and an output schema (which is not shown), the description should explain how to structure inputs. It doesn't describe the 'state' or 'questions' format beyond mentioning question types. While the output schema might clarify returns, the input side is underspecified, making the tool hard to use correctly without external knowledge. The advisory note adds safety context but doesn't fill the parameter gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'question types are choice, score, and noul', which hints at the structure of the 'questions' parameter but doesn't explain how to encode them or what 'state' should contain. Neither parameter is explicitly described, leaving an agent without enough information to construct valid inputs. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Run') and resource ('a local typed decision with Laya-MLX'), and enumerates the decision types (bounded classification, ordered scoring, yes/no probability). It doesn't explicitly differentiate from siblings (none exist), but the purpose is understandable and not a tautology. A slight vagueness in 'local typed decision' prevents a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use for bounded classification, ordered scoring, and yes/no probability', giving clear when-to-use guidance. It also includes a caution against using results for destructive actions, which is a form of when-not-to-use. Since there are no sibling tools, it can't name alternatives, but the guidance is otherwise explicit and practical.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedlaya_tell_me

TDQS

A3.6/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusing it with another. The tool's purpose is clearly distinct by default.

Naming Consistency4/5

The name 'laya_tell_me' is readable and uses a consistent imperative style with a product prefix. Since there is only one tool, there is no conflicting naming convention to penalize heavily.

Tool Count3/5

A single tool feels quite thin for a server, even if the intended scope is narrow. It is borderline and could benefit from at least one supporting tool, such as model inspection or configuration.

Completeness4/5

The tool appears to cover the advertised decision types: choice, score, and yes/no probability. Minor gaps like model management or detailed output explanations exist, but the core decision workflow is functional.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables running MCP tools against local MLX models on your Mac, with hardware-aware configuration, CLI streaming, and a dashboard for routing and monitoring.
    2,323 npm
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Laya-MLX as an MCP server (and plain-HTTP API). It answers typed decisions — choose an option / score a rubric / is this true? — on your own machine (Apple Silicon / MLX), with no cloud and ~10ms after warm-up. Not a chatbot. No token-by-token text, no JSON that can break. One forward pass returns a structured answer you can branch on. Ideal for routing, triage, classification, lead scoring, and g
    MIT