Skip to main content
Glama
bolohori
by bolohori

Windows Dev Agent

Windows-native developer orchestration for Claude Code and Codex, built around one shared runtime and six shared procedural skills.

Version 0.4.3 builds on 0.4.2 and closes host/runtime authority and outcome seams: Claude project-scoped tools are contained to the host-supplied project root, Codex plan shortcuts require authoritative session scope, Codex's MCP timeout covers the runtime execution ceiling, discovery failures remain canonical degraded snapshots, external-process audit outcomes preserve uncertainty, and WinGet installs are noninteractive and bound to the reviewed source.

Architecture

shared core
  skills/
  capabilities.yaml
  src/mcp/server.py
  src/capabilities.py
  src/discovery/
  src/models/environment.py
  src/observability/
  src/safety/gate.py
        │
        ├── Claude Code adapter
        │     .claude-plugin/plugin.json
        │     .mcp.json
        │     src/claude_server.py
        │     agents/  commands/  hooks/hooks.json
        │
        └── Codex adapter
              .codex-plugin/plugin.json
              .mcp.codex.json
              src/codex_server.py
              hooks/codex-hooks.json
              src/safety/codex_*.py
              src/observability/codex_*.py

Host adapters own only packaging, path binding, permission integration, and host hook contracts. Windows behavior and reusable procedures stay shared.

Related MCP server: PSKit

Shared skills

Skill

Purpose

env-inspect

Inspect Windows host, runtime, toolchain, package-manager, and isolation state without collapsing unknown probes into absence.

workflow-plan

Build a bounded Windows-aware plan when dependencies or uncertainty can change the route.

package-install

Resolve exact package identity, review the concrete mutation, execute under host permission, and verify fresh state.

sandbox-run

Select WSL, a project Dev Container, or Windows Sandbox from the required execution property.

win-setup

Repair missing or broken Windows development prerequisites with the smallest native change.

ecosystem-defrag

Inventory concrete agent/tool overlap and plan or execute a reversible consolidation.

Claude's /windows-dev-agent:env, /windows-dev-agent:plan, and /windows-dev-agent:defrag commands are thin adapters to those canonical skills. Codex consumes the same skills directly.

Install surfaces

Claude Code

From a checkout:

claude --plugin-dir .

.mcp.json keeps three identities separate:

  • ${CLAUDE_PLUGIN_ROOT} — immutable plugin code/config;

  • ${CLAUDE_PLUGIN_DATA} — persistent cache/audit data;

  • ${CLAUDE_PROJECT_DIR} — active project boundary.

Project-scoped Claude MCP calls are normalized to that host-supplied project root or one of its descendants. Relative project paths are interpreted under ${CLAUDE_PROJECT_DIR}; an arbitrary sibling/outside directory is rejected.

Codex / ChatGPT desktop

Canonical main carries .codex-plugin/plugin.json plus the release index at .agents/plugins/marketplace.json. Each published marketplace entry is pinned to an immutable Git commit rather than treating a moving main branch as the identity of a fixed version. The immutable plugin payload itself does not depend on carrying its own marketplace index.

Codex installs plugin code into its cache, so project-scoped MCP tools require the absolute current project directory explicitly. Runtime cache/audit state converges on ${CODEX_HOME:-~/.codex}/plugins/data/windows-dev-agent rather than the installed plugin tree.

Bundled Codex hooks are an optional trusted layer. Codex does not automatically trust newly installed/changed plugin hooks; until the user reviews and trusts them, native Codex MCP/shell approval policy remains the operative boundary. Mutation-capable WDA MCP tools remain prompt-gated regardless.

MCP surface

Both adapters expose the same ten tools:

Tool

Purpose

env_inspect

Native Windows environment snapshot with time-bound cache and force_refresh.

tool_discover

Focused executable/version discovery.

capability_run

Plan or execute one registered argv-based capability.

workflow_plan

Deterministic capability-aware scaffold that preserves weak, tied, unavailable, and selected route states.

package_search

Resolve package identity before mutation.

package_install

Plan or execute one exact WinGet/Chocolatey/Scoop install.

sandbox_run

Route and optionally launch WSL, Dev Container, or Windows Sandbox execution.

ecosystem_scan

Project-local agent/tool inventory by default; optional broader host inventory.

logs_query

Query bounded retained WDA audit metadata across every readable retained segment.

mcp_audit

Inspect project MCP configuration by default; broader reads are explicit.

Authority model

Windows Dev Agent does not carry a model-controlled user_approved token.

The sequence is deliberately simple:

model constructs exact tool call
→ active host applies its permission policy to that call
→ if permitted, the same call reaches the MCP runtime

Plan-only calls use execute: false. Mutation calls use execute: true. The shared runtime independently blocks capabilities classified forbidden; it does not pretend to prove that the host prompt occurred.

Claude Code

The Claude PreToolUse hook is tightening-only:

  • read-only / reversible → no WDA decision; Claude's normal policy remains authoritative;

  • approval-required / checkpoint → ask;

  • forbidden → deny;

  • it never returns allow.

Project-local ecosystem/MCP inspection is classified separately from broader host reads, but Claude Code's own normal permission flow remains authoritative for both.

Codex

Codex uses native MCP approval modes. Always-bounded non-filesystem reads can be approve; mutation-capable and filesystem-inventory tools remain prompt. The server-level MCP timeout is set above the runtime's 600-second execution ceiling so Codex does not terminate a legitimate long-running call before the runtime's own bound.

When Codex plugin hooks are trusted, PermissionRequest removes needless prompts only where the hook event itself proves the safe scope. package_install(execute:false) may be allowed as a non-executing plan. capability_run(execute:false) and workflow_plan may be allowed only when their absolute project path resolves inside the authoritative session cwd. sandbox_run planning remains on native approval because planning can enumerate payload paths. Executing calls receive no WDA allow decision and continue to normal Codex approval.

ecosystem_scan and mcp_audit remain on native Codex approval even for project-only requests. Their caller-supplied cwd is required, but the plugin does not auto-approve those filesystem reads.

PreToolUse may deny a known-forbidden action but otherwise defers prompting to native Codex policy.

Environment evidence

Availability fields are tri-state:

  • true — observed present/enabled;

  • false — observed absent/disabled;

  • null — the probe did not establish the fact.

The shipped PowerShell producer uses live-image optional-feature queries (Get-WindowsOptionalFeature -Online). Probe failures are recorded in snapshot.errors and make the snapshot degraded rather than silently becoming false. Windows Sandbox is queried using its canonical Containers-DisposableClientVM optional-feature identity. If the native discovery process times out or otherwise cannot produce a parsed snapshot, env_inspect still returns the same canonical snapshot schema with degraded/unknown state instead of switching to an ad-hoc failure shape.

The cache stores the same canonical representation returned to consumers; it is not a second serialization format. Cache TTL is five minutes, and package-install execution invalidates the cached snapshot even on a failed installer because partial mutation is possible.

Discovery intentionally does not persist username/domain, Git user identity, or full PowerShell module inventory. Those values are not required by the current routing consumers.

Package setup

package_search is the package-identity producer. It is non-mutating by intent, but it executes the selected package manager and may contact its configured source, so the active host remains authoritative for the call. A package mutation should use an exact ID from the user/authoritative project state or resolve one through search before package_install.

package_install(execute:false) returns the concrete argv for review. execute:true requests that exact mutation under the active host permission policy. WinGet execution is explicitly bound to the winget source and uses --disable-interactivity because MCP subprocess stdin is closed; package/source agreement flags are part of the reviewed argv. Installer exit is not treated as proof that the requested task now works; verify executable/version/task state afterward.

Workflow routing

workflow_plan separates semantic route evidence from execution availability. Capability ID/tag overlap supplies deterministic routing evidence; description-only similarity and generic operation words such as run, test, build, and lint do not select a capability by themselves. Minor token forms such as tests/test are normalized without turning generic words into discriminators. Equal strongest semantic matches remain ambiguous even when only one backend happens to be installed. A uniquely strongest semantic match with no executable is matched_unavailable; it does not silently fall through to a weaker match. Existing selection_status, matched_candidate, and selected_candidate fields remain stable, with route_discriminator added when the route is unresolved.

Isolation

environment:auto requires an isolation_requirement:

  • linux_compatibility → WSL;

  • project_reproducibility → configured project Dev Container;

  • untrusted_windows → Windows Sandbox.

WSL enters the active project using wsl --cd <project> and uses sh -lc by default. WSL accepts an absolute Windows path for --cd, so no guessed /mnt/<drive> translation is required. Dev Container execution uses the project configuration and sh -lc.

Windows Sandbox payloads

A hostile Windows workload is not isolated merely because a Sandbox window launches. For untrusted_windows, supply workspace-relative payload_paths identifying the files/directories the inner command actually needs. The runtime:

  1. rejects absolute paths, .. escapes, missing paths, symbolic links, NTFS reparse/junction boundaries, overlapping selections, payloads over 10,000 filesystem entries, and payloads over 1 GiB total file bytes;

  2. checks every selected path component before resolution, so a reparse parent cannot be crossed merely because its eventual target remains inside the workspace;

  3. fails closed when reparse metadata cannot be established;

  4. stages the selected payload into a temporary bundle;

  5. maps only that generated bundle into Windows Sandbox, read-only;

  6. disables Sandbox networking and clipboard;

  7. runs the inner command from C:\WDAShare\payload.

Planning does not materialize the bundle. An executing Sandbox call returns launched plus cleanup_path; that establishes launch only. Inner command success remains unknown until observed from inside the Sandbox. If staging or process launch fails before the Sandbox starts, the partial temporary bundle is removed automatically.

Ecosystem and MCP reads

ecosystem_scan starts project-local. Set include_host:true only when user-level extensions/plugins/MCP state can change the decision. include_packages:true is legal only with host inventory enabled.

mcp_audit likewise starts from the project boundary. User-level MCP configuration and arbitrary config_path reads are explicit broader requests.

Returned MCP summaries expose only structural metadata: server name, validity, transport kind, argument count, and whether command/URL/environment fields are present. Raw command strings, URLs, argument values, and environment values are not returned. JSON config reads are capped at 2 MiB per file. On Codex, filesystem inventory reads remain on native approval even when the supplied path is project-local.

Audit state and retention

WDA persists only metadata its audit consumers need; raw commands, MCP arguments, stdout/stderr, and arbitrary tool responses are not retained.

For execution-capable calls, the audit representation distinguishes:

  • succeeded — result establishes successful execution;

  • failed — result establishes failed execution;

  • unknown — execution may have started or occurred, but the available observation does not establish the effect;

  • not_executed — plan/block/unavailable/invalid input or a pre-launch failure prevented execution;

  • not_applicable — the lifecycle event has no execution outcome to classify, including permission/control events.

Hook lifecycle and execution outcome are deliberately separate. A PostToolUseFailure without a result does not prove that a potentially external action failed after starting; where the call could have launched a process, the execution/effect state remains unknown. Started calls that time out likewise remain unknown because partial or child-process mutation may already have occurred.

Codex PostToolUse may inspect the WDA MCP result in memory to derive that small status, then discards the raw response. A Windows Sandbox launch therefore remains unknown, never “zero failures.”

agent.log is bounded to 2 MiB and one rotated predecessor (agent.log.1). logs_query reads retained segments in chronological order and preserves other readable retained evidence if a listed segment disappears or becomes unreadable before acquisition. Environment cache and audit state live outside the immutable plugin code tree.

Capability catalog

capabilities.yaml contains only fields with live consumers: description, safety, tags, and argv tools. Configured commands execute with shell=False and DEVNULL stdin. Caller-supplied extra_args upgrade effective authority to approval-required rather than inheriting the base capability class.

The current catalog covers Git inspection, Python/JavaScript linting, Python/.NET tests, .NET build, and GitHub PR creation.

Verification

GitHub Actions runs on windows-latest across Python 3.9 and 3.13 and checks:

  1. Python runtime compilation;

  2. the contract-focused pytest suite, including real Windows junction containment;

  3. the actual shipped PowerShell discovery producer on the Windows runner;

  4. exact MCP initialization/version/tool surfaces for Claude and Codex;

  5. when the canonical release index is present, that its immutable SHA resolves to a byte-identical published payload outside the marketplace index itself.

The suite covers the one-call authority sequence, tri-state discovery/cache roundtrip, host/project read boundaries, argument-dependent safety, package freshness/source binding, routing discrimination, Sandbox payload staging/path/resource/reparse containment, WSL project binding, audit outcome uncertainty/retention, MCP summary minimization, host adapter wiring, and MCP transport isolation.

For local development:

python -m pip install -r requirements-dev.txt
python -m compileall -q src
python -m pytest tests -q

A green run supports only the contracts those checks actually exercise. It does not establish that a human accepted a real Claude/Codex permission dialog, nor that a command launched inside an interactive Windows Sandbox completed successfully on an end-user desktop.

Requirements

  • Windows 10/11 for full native behavior;

  • Windows PowerShell 5.1+;

  • Python 3.9+;

  • current Claude Code for the Claude adapter;

  • current Codex / ChatGPT desktop plugin runtime for the Codex adapter.

Runtime dependencies are Python standard-library only. Development tests use pytest.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    C
    maintenance
    Bridges Claude Desktop to local Windows 10 development environments to enable repository reading, documentation searching, and isolated code execution. It provides a secure interface for running allowlisted terminal commands and scripts directly through natural language.
    4
  • A
    license
    Not graded
    quality
    C
    maintenance
    Neural-safe PowerShell automation server for AI agents, enabling file management, git workflows, system inspection, and command execution with a 5-tier safety pipeline.
    MIT
  • A
    license
    A
    quality
    F
    maintenance
    A comprehensive Model Context Protocol (MCP) server that enables Claude and other LLM applications to execute PowerShell commands, scripts, and perform system operations on Windows systems.
    10
    25
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables AI assistants to execute PowerShell commands, manage files, inspect projects, run Git operations, and monitor system information on Windows through a local MCP server.

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/bolohori/Windows-Dev-Agent'

If you have feedback or need assistance with the MCP directory API, please join our Discord server