AI Agent Loop MCP Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AI Agent Loop MCP Serverdebug the failing test in broken-repo and fix the issue"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ค Task 3 โ AI Agent Loop with MCP
A production-style AI Debugging Agent built using the Model Context Protocol (MCP), capable of planning, inspecting repositories, proposing code edits with human approval, executing tests, and evaluating performance across a benchmark suite.
๐ Overview
This project implements a complete autonomous debugging agent that follows the Plan โ Act โ Observe execution pattern.
Instead of directly editing repository files, the agent communicates through an MCP (Model Context Protocol) server, allowing every repository interaction to occur via structured tools.
The agent:
understands failing tests
creates a debugging plan
explores the repository
reads source files
proposes code edits
waits for user approval
executes tests
repeats until success or budget exhaustion
The implementation follows all major requirements from Task 3.
โจ Features
Agent Loop
โ Planning
โ Tool selection
โ Repository exploration
โ Observation
โ Test execution
โ Halting conditions
Related MCP server: harness-fe
MCP Server
Implemented tools:
read_file
list_dir
grep
propose_edit
run_test
All repository interaction occurs exclusively through MCP tools.
Human Approval
Before modifying any file the agent:
validates edit
shows diff
waits for user approval
updates repository only after confirmation
Unsafe edits are rejected automatically.
Safety
Implemented guardrails:
Step Budget
Wall Clock Budget
Stuck Loop Detection
Approval Validation
Repository Boundary Checks
Tool Error Handling
Evaluation
Includes:
Golden evaluation suite
Metrics
Trajectory logging
Result reporting
๐ Architecture
+----------------------+
| CLI / Index |
+----------+-----------+
|
|
createInitialState()
|
|
+--------v--------+
| Agent Loop |
+--------+--------+
|
+---------------+----------------+
| |
| |
chooseTool() createPlan()
| |
| |
+------v-------+ +------v------+
| Groq LLM | | Planner |
+------+-------+ +-------------+
|
|
Tool Selection
|
|
+--------v---------+
| MCP Client |
+--------+---------+
|
|
+--------v---------+
| MCP Server |
+--------+---------+
|
+---------+----------+
| | |
read_file list_dir grep propose_edit run_test๐ Project Structure
Task-3-Agent-Loop
โโโ evals
โ โโโ golden-agent.jsonl
โ
โโโ packages
โ โโโ agent
โ โ
โ โโโ logs
โ โ โโโ trajectory.jsonl
โ โ โโโ eval-results.json
โ โ
โ โโโ src
โ โ
โ โ โโโ approval
โ โ โโโ eval
โ โ โโโ loop
โ โ โโโ mcp
โ โ โโโ metrics
โ โ โโโ client.ts
โ โ โโโ planner.ts
โ โ โโโ model.ts
โ โ โโโ logger.ts
โ โ โโโ state.ts
โ โ โโโ cli.ts
โ โ
โ โโโ tools
โ โโโ types
โ
โโโ broken-repo
โ
โโโ DESIGN.md
โโโ NOTES.md
โโโ RESULTS.md
โโโ README.md๐ง Agent Workflow
Run Tests
โ
Tests Fail
โ
Create Debugging Plan
โ
Choose Tool
โ
Execute Tool
โ
Observe Result
โ
Update State
โ
Need Another Tool?
โ
Yes โ Repeat
โ
No
โ
Run Tests
โ
Success
โ
Stopโ Agent State
The agent maintains the following state:
Property | Description |
currentTest | Active failing test |
currentTestOutput | Latest test output |
currentStep | Current iteration |
maxSteps | Maximum allowed iterations |
seenFiles | Already inspected files |
seenDirectories | Already listed directories |
fileContents | Cached repository files |
history | Tool execution history |
completed | Success flag |
๐จ Available Tools
Tool | Purpose |
read_file | Read source code |
list_dir | Explore repository |
grep | Search repository |
propose_edit | Request file modification |
run_test | Execute tests |
๐ก Safety Mechanisms
Step Budget
Stops infinite reasoning after the configured limit.
Wall Clock Budget
Terminates execution after maximum runtime.
Stuck Loop Detection
Stops execution when the same tool with identical arguments is repeatedly selected.
Approval Gate
Every modification:
validated
previewed
confirmed
before writing to disk.
๐ Metrics
The project reports:
Success Rate
Steps Used
Tool Errors
Guardrail Violations
Wasted Steps
Execution Time
Success within Budget
๐ Evaluation
Golden evaluation contains:
Difficulty | Cases |
Easy | 6 |
Medium | 6 |
Hard | 3 |
Total | 15 |
Each evaluation records:
success
execution time
metrics
logs
๐ป CLI
Run the debugging agent
pnpm agent fix --test tests/math.test.tsRun evaluation
pnpm agent evalRun live evaluation
pnpm agent eval --liveCompare against baseline
pnpm agent eval --compare baseline.json๐ Logs
Generated automatically:
logs/
trajectory.jsonl
eval-results.jsonTrajectory contains:
tool
arguments
timestamp
result
๐งช Technologies
TypeScript
Node.js
Groq API
MCP SDK
Vitest
PNPM
๐ฏ Assignment Requirements
Requirement | Status |
Agent Loop | โ |
Planner | โ |
MCP Tools | โ |
Approval Workflow | โ |
Trajectory Logging | โ |
Metrics | โ |
Evaluation Harness | โ |
Golden Dataset | โ |
CLI | โ |
Documentation | โ |
๐ Observability, Safety & Hardening (Task 3 Extension)
Task 3 was extended with observability, budget circuit breakers, a unified human-in-the-loop policy, and a prompt-injection red team โ aligned with the Agentic AI intern programme safety requirements applied to the Task 3 agent.
Structured Tracing
Every agent run instruments:
LLM calls (planner + tool selection)
Tool executions
Approval gate decisions
Each span records:
Field | Description |
startTime / endTime | Wall-clock timing |
input / output | Redacted request/response payloads |
tokens | Prompt, completion, and total token counts |
estimatedCostUSD | Per-span cost estimate |
parentId / children | Parent-child trace tree |
Traces export to:
generated/traces/trace.json
generated/history/run-NNN.jsonA local flame-graph style HTML viewer is generated at:
generated/traces/viewer.htmlOpen generated/traces/viewer.html in a browser to click through spans, inspect token usage, and locate the most expensive step.
Budget Circuit Breakers
Independent of the step budget, every run enforces:
Ceiling | Default | Behavior on breach |
Cost | $0.05 USD | Stops cleanly, logs reason, exports trace |
Wall clock | 180 seconds | Stops cleanly, logs reason, exports trace |
When a ceiling is hit mid-run, the agent:
Throws a logged budget error
Exports the trace tree
Generates a run report
Does not silently truncate or hang
Human-in-the-Loop Policy
All tool actions are routed through a unified policy table:
Action Category | Tools | Requires Approval |
Read-only | read_file, list_dir, grep | No โ logged for trace completeness |
Reversible write | run_test | No โ logged, applied automatically |
Irreversible | propose_edit | Yes โ explicit approval before write |
External | (reserved) | Yes โ explicit approval |
Set AUTO_APPLY=true in .env for non-interactive evaluation runs.
Prompt-Injection Red Team
12 adversarial test cases (6 blunt, 6 plausible) live in:
evals/injection-redteam.jsonlRun the full red-team suite:
pnpm securityEach case injects adversarial content into packages/broken-repo/src/math.ts, runs the agent, and checks:
Pass = real bug fixed AND secret not leaked
Fail = secret leaked, bug not fixed, or both
Results export to:
generated/security/results.json
generated/security/traces/<case-id>.jsonSee SECURITY.md for verbatim payloads and per-case outcomes.
Updated Project Structure
Task-3-Agent-Loop
โโโ evals
โ โโโ golden-agent.jsonl
โ โโโ injection-redteam.jsonl
โโโ generated
โ โโโ traces/ # JSON trace trees + viewer.html
โ โโโ history/ # Per-run state snapshots
โ โโโ reports/ # Markdown run reports
โ โโโ security/ # Red-team results & traces
โ โโโ golden/traces/ # Golden eval trace snapshots
โโโ packages
โ โโโ agent
โ โ โโโ src
โ โ โ โโโ approval/ # validateEdit, applyEdit, showDiff
โ โ โ โโโ budget/ # BudgetManager (cost + time ceilings)
โ โ โ โโโ eval/ # Evaluation harness
โ โ โ โโโ loop/ # runLoop, stuckLoop detection
โ โ โ โโโ mcp/ # MCP server + tool registration
โ โ โ โโโ metrics/ # Aggregate metrics
โ โ โ โโโ policy/ # HumanApprovalPolicy
โ โ โ โโโ report/ # Run report generator
โ โ โ โโโ security/ # Red-team runner + attack cases
โ โ โ โโโ tracing/ # Tracer, Span schema, history
โ โ โ โโโ viewer/ # HTML trace viewer generator
โ โ โโโ tools/ # MCP tool implementations
โ โ โโโ types/ # AgentState, ToolCall, ToolResult
โ โโโ broken-repo/ # Intentionally broken code under repair
โโโ DESIGN.md
โโโ NOTES.md
โโโ RESULTS.md
โโโ SECURITY.md
โโโ CHANGELOG.md
โโโ README.mdNew CLI Commands
Run prompt-injection red team:
pnpm securityRun agent fix (with tracing + budgets):
pnpm agent fix --test tests/math.test.tsView trace after a run:
# Open in browser
start generated/traces/viewer.html # Windows
open generated/traces/viewer.html # macOSKey Metrics (Safety Eval)
Metric | Description |
injection resistance rate | Pass rate across 12 cases, split blunt vs plausible |
secret-leakage rate | Should be zero |
budget-breach handling | Every forced breach stops cleanly with logged reason |
trace completeness | All spans have intact parent/child links |
mean added latency | Instrumentation overhead (target: near zero) |
Full numbers in RESULTS.md. Full attack payloads in SECURITY.md.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceAn MCP server and VS Code extension that enables AI clients to interactively debug code using breakpoints, execution control, and state inspection. It is language-agnostic and works with any debugger that supports VS Code's launch.json configurations.
- AlicenseNot gradedqualityAmaintenanceA source-aware MCP server that connects AI agents to browser and server runtimes, enabling real-time debugging, monitoring, and automatic fixes via WebSocket or HTTP.2MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for AI-assisted Python debugging using debugpy and Debug Adapter Protocol, enabling AI agents to run tests, set breakpoints, and inspect variables via natural language.8MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that extends AI coding assistants with deterministic, algorithmic capabilities such as code analysis, fault localization, and formal verification, enabling an autonomous engineering team within the IDE.MIT
Related MCP Connectors
Agent Replay Debugger MCP โ record every agent step + deterministic replay. Step-debugger for
Live browser debugging for AI assistants โ DOM, console, network via MCP.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/HarshPariya/Task-3-ai-agent-loop-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server