Forgehand
Forgehand is an MCP server that lets a supervisor (like Codex) delegate bounded coding tasks to a local OpenAI-compatible model inside isolated Git worktrees.
Check health and configuration:
forgehand_healthreturns the active worker, safety limits, registered repositories, and available paths.Delegate implementation tasks:
forgehand_delegateruns a complete task with an objective, scoped file paths, acceptance criteria, optional constraints, commands, and required command IDs.Control task execution: Set edit mode, max iterations, base revision, whether changes must be present, keep the worktree, allowed actions, and explicit risk acknowledgement for host commands.
List recent tasks:
forgehand_tasksreturns compact receipts for recent tasks without loading raw logs or diffs.Fetch a specific result:
forgehand_resultreturns one compact task receipt by its task ID, including patch, tests, token usage, and uncertainties.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ForgehandFix the failing tests in src/parser using a local model and show me the diff."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Forgehand is a small, model-agnostic MCP server for delegating complete coding tasks to an OpenAI-compatible local model. Each step receives bounded current state, deterministic repository facts, and only the latest observation — never the accumulated chat transcript.
Bounded batch reads let the worker inspect up to eight scoped files in one model round-trip without turning the conversation transcript into memory.
Why
Cloud coding agents are excellent architects and reviewers. They are also an expensive place to repeat file reads, compiler runs, failed tests, and mechanical edits. Forgehand moves that inner loop to your machine and returns a compact receipt with the patch, tests, token usage, and uncertainties.
You → Codex → one task contract → Forgehand → one local model
↑ │
└──── compact receipt ────┘Related MCP server: cc-delegate
Quick start
Requirements: Python 3.11+, Git, Codex, and one local server exposing an
OpenAI-compatible /v1/chat/completions endpoint with structured JSON output.
python -m pip install "forgehand-local @ git+https://github.com/JVCSampaio/forgehand.git"
forgehand init
forgehand repo add /path/to/your/repository
forgehand doctor
codex mcp add forgehand -- forgehand mcpThe default example is Ornith 1.5 9B at http://127.0.0.1:1234/v1. Choose any
compatible local model during setup:
forgehand init --model your-model-id --worker-name "My local worker" --forceThen ask Codex:
Use Forgehand to update the parser. Only modify
src/parserandtests/parser. Require the approved parser tests to pass, inspect the final diff, and review the result. I acknowledge that approved commands inherit host permissions and network.
Codex can call four MCP tools: forgehand_health, forgehand_delegate,
forgehand_tasks, and forgehand_result.
Open the private, read-only dashboard at any time:
forgehand dashboardIt binds only to 127.0.0.1, hides full repository paths, and reports measured
local-worker usage. It never presents those numbers as estimated Codex savings.
What stays bounded
exact Git roots must be registered by the user;
every task runs in a detached worktree;
the worker can access only declared repository-relative paths;
commands are supervisor-provided
argvarrays selected by ID, without a shell;host commands require an explicit risk acknowledgement in every task contract;
required command IDs must all exit zero before
successis accepted;common credential-bearing environment variables and interactive Git prompts are removed;
state, observation, output, steps, changed files, and retries have hard limits;
repeated commands on an unchanged tree are rejected, and action-rejection budgets stop unproductive local loops early;
implementation contracts can set
requires_changes=trueso an empty diff cannot be reported as success;full logs and diffs stay local; compact receipts return to the supervisor;
Codex review is always required before integration.
Forgehand does not provide an OS-level process sandbox. Approved commands still inherit the host network stack and OS permissions, even though they run without a shell and with a scrubbed environment. Only approve commands you trust. See the task contract and security model.
Token evidence
In one same-contract microbenchmark, the original conversational worker loop used 25,826 local-worker tokens. Forgehand used 10,708 — 58.54% fewer — while producing the same one-file change. Forgehand was 5.52% slower in that run because the worker needed several structured-output retries.
This is preliminary local-worker evidence, not a promise of 58.54% Codex, API, credit, cost, or latency savings. See Token accounting.
Design
Forgehand is inspired by the bounded-state execution described in SKILL.state. It adds coding-specific controls: transactional SQLite revisions, typed state operations, deterministic Git facts, scoped tools, worktree isolation, durable artifacts, and per-attempt usage.
History is archived, not attended.
Forgehand deliberately ships without self-evolving skills, multi-model routing, an unrestricted shell, or automatic cloud fallback. Experimental learning belongs in a separate optional project until it demonstrates net savings.
Status
Forgehand is alpha software. Use it on repositories with version control, inspect every patch, and rerun critical validation yourself.
License and attribution
MIT licensed. Forgehand evolved from ideas and MIT-licensed work in RANJIANG23/codex-hermes-worker. See NOTICE for attribution.
Available Tools
4 toolsforgehand_delegateA
Run one complete implementation task with bounded files and commands.
The repository must be registered first. Scope contains the only paths the worker may read or edit. Commands are supervisor-authored argv arrays invoked by ID without a shell, but inherit host permissions and network. Set the risk acknowledgement explicitly whenever commands are present. Required command IDs must all exit zero before the worker can return success.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | Yes | ||
| commands | No | ||
| edit_mode | No | free | |
| objective | Yes | ||
| constraints | No | ||
| base_revision | No | HEAD | |
| keep_worktree | No | ||
| max_iterations | No | ||
| allowed_actions | No | ||
| repository_root | Yes | ||
| requires_changes | No | ||
| acceptance_criteria | Yes | ||
| required_command_ids | No | ||
| acknowledge_host_command_risk | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by explaining that commands inherit host permissions and network, risk acknowledgement is required when commands are present, required command IDs must all exit zero for success, and the worker is bound to the scope. Annotations carry destructiveHint=false and readOnlyHint=false, and the description adds important context about the command execution model without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the primary purpose in the first sentence, and every sentence adds operational detail without redundancy. It covers command execution semantics, safety, and success conditions in about 60 words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (14 parameters, command execution, security considerations) and the presence of an output schema, the description covers the most critical operational context. It doesn't explain every parameter but the essentials for safe and correct invocation — registration prerequisite, scope limitations, command execution model, and success criteria — are present. An output schema exists, so return values need not be explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameter semantics. It explains the significance of scope, commands, risk acknowledgement, and required command IDs, which maps to key parameters. Not every parameter is individually explained (edit_mode, base_revision, etc.), but the core ones an agent needs to invoke correctly are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs 'one complete implementation task' with bounded files and commands, distinguishing it from sibling tools like forgehand_health, forgehand_tasks, and forgehand_result which are inspection/retrieval tools. The verb 'Run' plus the resource and scope constraint makes the purpose explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for use: repository must be registered first, scope bounds what the worker may access, and commands are supervisor-authored. It doesn't explicitly name sibling alternatives or when not to use this tool, but the distinction from the health/tasks/result siblings is implied by the delegation purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forgehand_healthARead-onlyIdempotent
Return active worker, safety limits, registered repositories, and paths.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds value by specifying the exact information returned (active worker, safety limits, registered repositories, paths), providing concrete context beyond the annotations. It does not contradict any annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the action, and lists all key output categories with no filler or redundancy. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only health check with an output schema present, the description is fully sufficient. It names all the key information an agent would need to evaluate the tool's relevance and expected result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter semantics to enrich. The description correctly focuses on the output rather than inputs, and the baseline for zero-parameter tools is 4. No additional parameter explanation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and a clear resource: active worker, safety limits, registered repositories, and paths. This strongly differentiates it from siblings like forgehand_delegate, forgehand_tasks, and forgehand_result, which appear to be action-oriented operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a health/status-check use case but does not explicitly state when to use this tool versus the sibling tools, nor does it mention any alternatives or exclusions. The purpose is inferable from the name and return value list, but no direct guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forgehand_resultARead-onlyIdempotent
Return one compact task receipt by ID.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal readOnly, idempotent, and non-destructive behavior, and the description's 'Return' is consistent with those. It adds the nuance of 'compact' receipt but does not disclose non-obvious behaviors such as error handling, whether missing IDs return null/error, or any polling semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is an eight-word single sentence that states the core operation immediately. There is no filler, redundant restatement of the tool name, or optional background that could be dropped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter read with rich annotations and an output schema, this is minimally viable. However, it stope short of explaining how the task_id relates to forgehand_delegate or forgehand_tasks, and whether a receipt exists immediately after delegation or after completion, leaving contextual gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description shoulders the burden of explaining task_id, but 'by ID' adds little beyond the parameter's own name and title 'Task Id'. It does not describe the ID format, where to obtain it, or any constraints, so parameter semantics remain under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb ('Return') and noun phrase ('one compact task receipt') with a clear qualifier ('by ID'), which plainly identifies what the tool does. It distinguishes itself from the sibling tools like forgehand_tasks (likely listing) and forgehand_health by focusing on a single item lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'by ID' implies it should be used when the caller already has a task_id and wants a single receipt, but the description does not explicitly say when to use this over forgehand_tasks or forgehand_delegate. No alternatives or exclusion criteria are given, leaving usage to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forgehand_tasksARead-onlyIdempotent
List compact receipts for recent tasks without loading raw logs or diffs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds useful behavioral context by stating it returns compact receipts and deliberately avoids loading raw logs or diffs, clarifying the tool's lightweight nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise, front-loaded sentence that communicates the action, resource, and key behavioral qualifier without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter list tool, the description, annotations, and output schema provide enough context for correct invocation. The main gap is explicit parameter semantics, but overall completeness is solid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'limit' parameter has no schema description, and schema description coverage is 0%. The tool description does not explain what 'limit' controls or how it affects results, so it fails to compensate for the missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List compact receipts for recent tasks'. The qualifier 'without loading raw logs or diffs' clarifies the output format and distinguishes it from heavier task-detail tooling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys a clear context: use this for a lightweight overview of recent tasks, and not when raw logs or diffs are needed. It does not explicitly name sibling alternatives or exclusion conditions, but the intended use is evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: health reports status, delegate starts execution, tasks lists receipts, and result retrieves one receipt by ID. There is no meaningful overlap or ambiguity between them.
All tools use the same forgehand_ prefix and snake_case style, but the second segment is inconsistent in form: delegate is a verb while health, tasks, and result are nouns. The pattern is still predictable and readable.
Four tools is well-scoped for a focused task-execution server. The set covers status, submission, listing, and retrieval without unnecessary extras or obvious bloat.
The core submit-to-retrieve workflow is covered along with health/status, but there is no cancellation tool, and repository registration appears to be a prerequisite handled outside the MCP surface. These are minor gaps for a one-shot task runner.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Persistent memory and cross-session learning for AI coding assistants (hosted remote MCP).
AI-native git hosting — repos, PRs, issues, CI gates, and AI code review over MCP (60 tools).
Share one project context across ChatGPT, Claude, Telegram and any MCP client.
Related MCP Servers
- AlicenseAqualityBmaintenanceDelegate coding tasks to external AI coding agents in isolated git worktrees with independent verification, enabling any MCP client to orchestrate multi-agent workflows.5MIT
- AlicenseAqualityAmaintenanceDelegates heavy development tasks from a supervisor to an autonomous worker on a cheaper model via MCP, with isolated git worktrees and provider-agnostic support.31MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that lets ChatGPT or any MCP client securely delegate coding tasks to a local Claude Code instance, with git checkpointing, approval gates, and structured results. Supports code review, test running, and rollback via simple tool calls.16MIT
- AlicenseAqualityAmaintenanceA local MCP server that delegates coding tasks to a temporary OpenCode session and returns a completion report with changed files, tool calls, cost, and the subagent's reply.15181MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/JVCSampaio/forgehand'
If you have feedback or need assistance with the MCP directory API, please join our Discord server