MCP WebPilot
Connects to a local multimodal Ollama model (default gemma3:4b) that powers the agent's reasoning, using both text reasoning and browser screenshot understanding to decide one action at a time within the observe → decide → act loop.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP WebPilotgo to example.com and tell me the main heading"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP WebPilot
B.Tech Final-Year Project — Local Multimodal MCP Browser Agent
MCP WebPilot connects a local multimodal Ollama model to browser automation through the Model Context Protocol (MCP). The agent observes the current webpage using both structured DOM information and a browser screenshot, chooses one safe action, executes it through Playwright, and repeats until the task is complete or a safety boundary is reached.
Architecture
User
│
▼
Frontend
│
▼
FastAPI Agent Loop
│
├── Ollama (Gemma 3 4B)
│ ├── text reasoning
│ └── image understanding
│
├── Safety / Audit
│
▼
MCP Server
│
▼
Playwright
│
▼
Chromium
├── DOM/accessibility state
└── screenshot
│
└──────► next agent decisionFeatures
MCP Streamable HTTP server
FastAPI REST API
Playwright Chromium automation
Local Ollama multimodal model
Iterative observe → decide → act → observe loop
Screenshot-aware agent decisions
DOM element inventory and robust locator fallbacks
Persistent browser profile
Safety gate and audit log
Human-controlled boundary for irreversible workflows
Live browser screenshot in the UI
Docker + Compose
Lightweight vanilla frontend
Pytest + GitHub Actions
Ollama model
The default model is gemma3:4b for both text and image input. Using one multimodal model avoids downloading a separate vision model.
ollama pull gemma3:4b
ollama run gemma3:4bThe model remains outside the Git repository.
Run locally
python -m venv .venv
# Windows
.venv\\Scripts\\activate
# macOS/Linux
# source .venv/bin/activate
pip install -r requirements.txt
python -m playwright install chromium
uvicorn backend.main:app --reload --port 8000Open http://localhost:8000.
Docker
docker compose up --buildOpen http://localhost:8000.
For Docker, Ollama normally runs on the host and is reached through host.docker.internal.
Agent workflow
Example:
User goal
↓
Observe DOM + screenshot
↓
Ollama chooses one action
↓
Backend validates action
↓
Playwright executes it
↓
Observe new state
↓
RepeatThe model is an untrusted planner. The backend is always the execution authority.
Safety
The agent can navigate, inspect pages, click controls, and fill ordinary fields. Sensitive irreversible actions are stopped at the safety boundary and require explicit human handling.
The system does not provide unrestricted automation for transactions, security bypasses, credential theft, or verification challenges.
Storage target
The repository does not contain models, browser binaries, virtual environments, datasets, or caches. The selected Ollama model is about 3.3 GB according to the current Ollama model listing, while project/runtime dependencies are kept lightweight. Exact total disk usage depends on the operating system, Python environment, browser cache, and Ollama storage.
Project structure
frontend/ Static UI
backend/main.py FastAPI application
backend/mcp_server.py MCP tool server
backend/services/ Browser + Ollama + agent
backend/utils/ Safety + audit
tests/ Automated tests
docs/ Architecture and safety
Dockerfile Container image
compose.yaml Docker Compose
requirements.txt Python dependenciesAcademic scope
This project demonstrates MCP protocol integration, multimodal local-LLM reasoning, browser automation, agent loops, API design, safety engineering, containerization, testing, and auditability for a B.Tech final-year project.
This server cannot be deployed
Maintenance
Related MCP Connectors
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Real Chrome for agents: start a browser, read pages as numbered markdown, click, type, hand off.
Supervised API-write gateway for AI agents with policy, human approval and execution receipts.