CodeForgeX
CodeForgeX: Deterministic AI-Agent Evaluation Environment & MCP Software Engineering Harness
CodeForgeX is an enterprise-grade, deterministic AI-agent software engineering evaluation environment and tool-calling execution harness. Built on the official Model Context Protocol (MCP Python SDK v2), it provides an isolated, uncheatable sandbox where AI agents explore repositories, reproduce failures, formulate hypotheses, apply unified diff patches, and verify solutions against public and hidden test suites.
Unlike subjective "LLM-as-a-judge" grading, CodeForgeX employs deterministic, multi-criteria verification with cryptographic test tampering detection, regression guards, and wall-clock execution limits.
Architecture Topology
flowchart TD
subgraph ControlPlane["CLI & Orchestration (scripts/run_evaluation.py)"]
CLI["CLI Runner"] --> Runner["EvaluationRunner"]
CLI --> Loop["AgentExecutionLoop"]
end
subgraph AgentLayer["Autonomous Agent & Planning (src/agent/)"]
Loop --> Planner["SystematicSWEPlanner / LLMPlanner"]
Planner -->|AgentAction| Loop
Loop -->|call_tool| Client["MCPClient (Async stdio)"]
end
subgraph ProtocolBoundary["Model Context Protocol Boundary"]
Client <==>|Anonymous OS Pipes (stdio JSON-RPC)| Server["MCPServer (Subprocess)"]
end
subgraph ToolSurface["Sandboxed Tool Surface (src/mcp_server/)"]
Server --> T1["list_files"]
Server --> T2["read_file"]
Server --> T3["search_code"]
Server --> T4["apply_patch"]
Server --> T5["get_git_diff"]
Server --> T6["run_tests"]
Server --> T7["get_test_output"]
Server --> T8["get_repository_status"]
end
subgraph SecurityBoundary["Security & Anti-Cheat Sandbox (src/mcp_server/security/)"]
T1 & T2 & T3 & T4 & T5 & T6 & T7 & T8 --> Sandbox["SecurityPolicy Engine"]
Sandbox --> Guard1["Path Traversal Containment"]
Sandbox --> Guard2["Anti-Cheat (Hidden Tests Isolated)"]
Sandbox --> Guard3["Command Whitelisting & Injection Defense"]
end
subgraph TaskWorkspaces["Ephemeral Sandboxes (tasks/)"]
Guard1 & Guard2 & Guard3 --> TargetRepo["Ephemeral Workspace (.git)"]
end
subgraph Evaluator["Deterministic Evaluator Engine (src/evaluator/)"]
TargetRepo --> Verifier["TaskVerifier"]
Verifier --> ShaCheck["SHA-256 Digest Tamper Check"]
Verifier --> PublicRun["Public Test Execution"]
Verifier --> HiddenRun["Privileged Hidden Test Suite"]
Verifier --> DiffAnalysis["Git Diff & Line Metrics"]
ShaCheck & PublicRun & HiddenRun & DiffAnalysis --> Scorer["ScoringEngine (100 Pt Model)"]
Scorer --> Artifacts["Telemetry Artifacts (results/eval_*.json, .md)"]
endKey System Capabilities
Official Model Context Protocol (SDK v2): Real-time tool discovery and execution over standard I/O (
stdio), eliminating TCP port exhaustion and network race conditions.Multi-Provider Schema Reflection: Dynamic tool schema conversion supporting OpenAI function calling, Anthropic Claude, and Google Gemini function declarations.
Multi-Dimensional Scoring (0 - 100 Points):
Task Completion (40 pts)
Hidden Verification Tests (25 pts)
Public Baseline Tests (15 pts)
Regression Safety (10 pts)
Tool Efficiency & Token Conservation (5 pts)
Patch Conciseness & Quality (5 pts)
Anti-Cheat & Anti-Tampering Protection: Dual-layer defense with cryptographic SHA-256 test file digest verification. Modifying or deleting test assertions forces an immediate
0.0 / 100.0disqualification score.Systematic SWE Reasoning Workflow: Enforces Test-Driven Software Engineering:
EXPLORE$\rightarrow$REPRODUCE$\rightarrow$ANALYZE$\rightarrow$PATCH$\rightarrow$VERIFY$\rightarrow$FINISH.Fault-Tolerant Circuit Breakers: Active consecutive-failure guards prevent runaway token expenditure and model hallucination loops.
Hardened Docker Isolation: Non-root user execution (
uid=1000),cap_drop: [ALL],no-new-privileges:true, and RAM-backed in-memorytmpfsmounts.Unified Benchmark Dashboard: Terminal dashboard with JSON and Markdown artifact generation.
Benchmark Suite Catalog
CodeForgeX includes 6 diverse benchmark tasks spanning core software engineering modalities:
Task ID | Category | Difficulty | Problem Domain | Baseline Defect | Golden Score |
|
| Easy | E-Commerce Pricing Engine | Flat rate subtraction instead of % calculation; no bound checks |
|
|
| Medium | Concurrent LRU Cache with TTL | Expired nodes not purged on access; evicts MRU instead of LRU |
|
|
| Medium | Thread-Safe Token Bucket Limiter | Class raises |
|
|
| Medium | Request Dispatcher to Strategy | Monolithic |
|
|
| Medium | Log Stream Deduplication ($O(N^2) \rightarrow O(N)$) | $O(N \times W)$ quadratic nested search takes > 1.5s and times out |
|
|
| Medium | Topological Build Dependency Sorter | Lacks topological sort and Tarjan/DFS cycle path detection |
|
Quickstart Guide
1. Installation & Environment Setup
# Clone the repository
git clone https://github.com/kirubesh/CodeForgeX.git
cd CodeForgeX
# Install virtual environment and dependencies using uv or pip
pip install -e .2. Run the Full Test Suite (91 Tests)
pytest -v3. Run Benchmark Evaluations via CLI
# Run the entire benchmark suite in golden reference mode
python scripts/run_evaluation.py --all
# Run a specific task in autonomous agent mode (MCP stdio tool-calling loop)
python scripts/run_evaluation.py --task bug_fix_001 --mode agent-systematic
# Filter benchmarks by category or difficulty
python scripts/run_evaluation.py --category algorithm
python scripts/run_evaluation.py --difficulty medium4. Containerized Execution with Docker
# Build the security-hardened container
docker build -t codeforge-x:latest .
# Run full evaluation across all tasks inside Docker
docker run --rm codeforge-x:latest --evaluate-all
# Or run using Docker Compose with cgroups limits and tmpfs in-memory sandboxes
docker compose up codeforge-evalTechnical Documentation & Interview Defenses
System Architecture Specification: Deep-dive into MCP server tools, security sandboxing, scoring math, and agent loop lifecycle.
Security & Threat Modeling: Comprehensive threat matrix covering path traversal, shell injection, privilege elevation, and anti-cheat guards.
Docker Reproducibility Guide: Container security specs, capability dropping, and cgroup resource quotas.
Staff-Level Technical Interview Defense Guide: Complete question-and-answer handbook defending every architectural trade-off and systems engineering choice.
License
MIT License. See LICENSE for details.