ForgeMCP
Provides integration with GitHub repositories through MCP, enabling permissioned read and write operations on repository content and GitHub resources.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ForgeMCPFix the parser regression and add a focused test"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ForgeMCP — Repository-Level Coding Agent
ForgeMCP accepts a software issue, gathers repository evidence, edits code through policy-checked tools, and refuses to declare success until tests verify the patch.
It is intentionally not a collection of MCP integrations. The project isolates three runtime questions that determine whether a coding agent can solve multi-file work reliably:
How should the runtime close the loop between planning, action, and verification?
Which repository evidence earns a place in a fixed context budget as code changes?
How can tools remain useful without becoming an unrestricted shell?
Stack: Python · MCP · Tree-sitter · OpenAI Responses API · Docker · Pytest · SQLite
Auditable benchmark snapshot
The checked-in formal benchmark uses eight synthetic repository defects, three paired
trials, the pinned gpt-5.2-2025-12-11 snapshot, identical 12-model-call budgets, and
external graders. Each arm therefore contains 24 attempts with raw per-run records.
Formal-8 aggregate | Baseline | Hybrid | Hybrid change |
Strict solve rate | 17/24 (70.8%) | 17/24 (70.8%) | +0.0 pp |
Average tool calls | 9.42 | 7.79 | −17.3% |
Average input tokens | 22,258 | 16,521 | −25.8% |
Repeated-read ratio | 30.2% | 25.6% | −4.6 pp |
Estimated cost per solve | $0.0649 | $0.0506 | −22.0% |
All 48 patches passed their hidden graders and no run changed a protected test. The
strict score is lower because seven runs per arm exhausted the 12-call budget before a
succeeded terminal state. See the
formal report, its six raw reports, and the
protocol for hashes and limitations.
The original local Git object database and 120-instance raw records were lost. The
historical headline summary remains in
benchmarks/reconstructed-swebench-verified-120.summary.json
as provenance only; it is not presented as a reproduced result.
Related MCP server: MCP RepoBridge
See it work in under a minute
git clone https://github.com/Tz22z/ForgeMCP.git
cd ForgeMCP
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
forge demoThe offline demo starts from a real two-file checkout bug, runs the same runtime with a
recorded model transcript, edits shop/pricing.py and shop/checkout.py, executes
independent assertions, and prints the verified diff. It needs no API key.
succeeded
Files changed 2
Tool calls 5
Model calls 6
Tests passedSee the demo walkthrough for the issue, expected trace, and interview talk track.
Runtime architecture
flowchart LR
I[Issue] --> S[Task state machine]
S --> X[Incremental index]
X --> C[Budgeted context selector]
C --> M[Model adapter]
M --> D{Strict tool dispatcher}
D --> F[File tools]
D --> G[Git inspection]
D --> T[Test runner]
F --> X
T -->|failure + locations| C
T -->|passing evidence| V[Verified patch]
B[Call · token · time · repeat budgets] --> S
P[Path · schema · approval policy] --> D
J[Append-only event journal] --- S
J --- DVerified execution loop
A typed state machine makes
planning,acting,verifying, and terminal outcomes observable rather than implicit in prompt text.Tool-call, model-call, input-token, output-token, wall-time, and repeated-action limits are checked centrally before work proceeds.
Every edit advances a mutation generation. Completion is valid only if the current generation has a passing test result.
Failed verification is compacted, its file/line evidence is fed back into retrieval, and the model gets another bounded repair turn.
Exhaustion returns
budget_exhausted; it is never silently converted into success.
Repository context that follows the code
A persistent SQLite index stores content hashes, Tree-sitter symbols, imports, parser provenance, and read counters.
Refreshes reparse only changed files and remove deleted entries.
An edit invalidates the changed file plus dependency consumers before the next model turn.
Ranking combines issue/path terms, symbols, dependents, recent Git diffs, test pairs, and traceback locations.
Snippets are packed under a hard estimated-token budget. Long tool output keeps a bounded head/tail plus a durable reference to the full observation.
This follows the broader context-engineering principle of treating context as a finite attention budget and favoring compact, just-in-time evidence over exhaustive dumps; see Anthropic's context engineering discussion.
Restricted tools
ForgeMCP exposes nine deliberately narrow tools: file listing, bounded reads, search, exact replacement, explicit file creation/overwrite, content-hashed deletion, Git diff, Git status, and tests.
Strict Pydantic schemas reject unknown arguments.
All paths resolve beneath one repository root; absolute paths and traversal fail.
.git,.forgemcp, and environment files are protected from model writes.Test arguments use an allowlist, subprocesses receive argument vectors without a shell, and high-risk operations enter the approval path.
Docker mode blocks network by default, drops every capability, sets
no-new-privileges, uses a read-only root filesystem, and caps CPU, memory, PIDs, and temporary storage.
The fuller threat model is in SECURITY.md.
Live usage
Create an issue file:
# Cache entries survive invalidation
Deleting a project must invalidate both the project cache and its derived permission
cache. Add a regression test and preserve the public API.Then run:
export OPENAI_API_KEY=...
forge index /path/to/repository
forge context "cache entries survive invalidation" --repo /path/to/repository --tokens 6000
forge run issue.md --repo /path/to/repository --model gpt-5.2 --dockerThe OpenAI adapter uses the Responses API with strict custom function definitions,
function-call outputs, and previous_response_id, following the
official OpenAI Responses API reference.
MCP server
cp mcp.json.example mcp.json
# Set FORGEMCP_ROOT in the copied file, then configure your MCP client.
forge serveThe server exposes:
MCP tool | Purpose | Mutates repository |
| Incrementally refresh symbols and dependencies | No |
| Preview ranked snippets and token use | No |
| Run the full edit-and-test loop | Yes |
| Read a compact event timeline and final result | No |
MCP repository arguments are relative to FORGEMCP_ROOT; absolute paths and escapes
are rejected at the server boundary.
Configuration
Copy .forgemcp.toml.example to .forgemcp.toml in a target repository. Important
settings:
[agent]
model = "gpt-5.2"
context_strategy = "hybrid" # or "baseline" for the ablation
context_token_budget = 12000
[budget]
max_tool_calls = 30
max_model_calls = 12
max_input_tokens = 80000
max_output_tokens = 20000
max_wall_seconds = 900
max_repeated_actions = 2
[tools]
approval_mode = "on-risk"
approved_tools = [] # Add "delete_file" only after explicit review
execution_mode = "docker"
allow_network = false
test_command = ["python", "-m", "pytest", "-q"]Reproduce an evaluation
Evaluation graders live outside the agent workspace. The harness also hashes visible tests before and after each run, so changing either public or hidden scoring conditions cannot count as a solve.
forge evaluate benchmarks/your-manifest.json \
--strategy baseline \
--model gpt-5.2 \
--output artifacts/baseline.json
forge evaluate benchmarks/your-manifest.json \
--strategy hybrid \
--model gpt-5.2 \
--output artifacts/hybrid.jsonPass explicit per-million-token prices if cost estimates are needed; ForgeMCP does not hard-code volatile model pricing. Reports include solve rate, tool/model calls, token usage, estimated cost per solve, repeated-read ratio, grader status, and changed files. See evaluation protocol.
For a low-cost live integration check, use the four-case pilot. For the current auditable result, use the formal eight-case suite, which checks in its manifest, external graders, three paired trials, raw records, and input/report hashes. Neither synthetic suite is presented as SWE-bench.
Development
make quality # lint + formatting + tests + strict mypy
make demo # deterministic end-to-end repair
make image # build the runtime containerThe current suite covers budgets, transitions, event redaction, scanners, symbol extraction, incremental invalidation, context packing, tool policy, output truncation, worktree isolation, Docker flags, model adaptation, automatic verification, MCP root confinement, and hidden-test grading.
Project boundaries
This is a research-quality agent runtime, not a promise that arbitrary generated code is safe to execute on a production host.
The token estimator is deliberately provider-independent and conservative; provider usage returned by the API remains the source of truth for accounting.
Tree-sitter falls back to Python's AST when a parser is unavailable, keeping offline runs functional.
Historical benchmark percentages are reconstructed summaries. Fresh claims should ship with immutable manifests, raw per-instance records, model IDs, and pricing inputs.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
AI-native git hosting — repos, PRs, issues, CI gates, and AI code review over MCP (60 tools).
Coordinate coding agents through MCP using existing AI plans, saved work, and independent checks.
Project management MCP for AI agents with safe task reads and writes.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEmpower any MCP-compatible AI Agent(MCP Client) with engineering-grade capabilities to understand, modify, run, and deliver real-world code repositories.348 PyPI1,155Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables safe repository inspection and Docker-sandboxed command execution via MCP, with optional Codex handoff for implementation tasks.1Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables AI coding agents and hosts to enforce deterministic repository boundaries via MCP, providing structured reads, supervised edits, snapshots, audits, and recovery with machine-readable evidence.MIT
- AlicenseNot gradedqualityAmaintenanceEnables repository-aware lifecycle management and controlled delegation of coding tasks to trusted worker harnesses via MCP, with execution isolation, recovery, and verified handoff.6 npm8Apache 2.0